Submind YouTube summaries
Thumbnail for SIB Bioinformatics Award ceremony 2025 - Early Career Award - Michael Skinnider

SIB Bioinformatics Award ceremony 2025 - Early Career Award - Michael Skinnider

Watch on YouTube

Video summary

Michael Skinnider, the recipient of the 2025 Early Career Award from SIB Bioinformatics, has established himself as a distinguished computational biologist with significant achievements across diverse fields including genomics, transcriptomics, proteomics, and metabolomics. His career trajectory began at Hamilton University in Canada, where he developed tools to discover bacterial secondary metabolites, before expanding his research into protein interaction networks and single-cell analysis during his MD-PhD studies. Notably, his work on identifying specific subpopulations of neurons contributed to breakthroughs in restoring walking ability for paralyzed patients through electrical stimulation of the spinal cord. His versatility was further highlighted by his successful defense of the MD-PhD program in Canada and his immediate appointment as a faculty member at Princeton University following his doctoral studies. The core of Skinnider's award-winning research addresses a fundamental gap in biological science: while scientists can routinely measure entire genomes, transcriptomes, and proteomes, comprehensively identifying small molecules in metabolomics remains a formidable challenge. He describes this unidentified portion of the metabolome as "chemical dark matter," which his lab aims to illuminate using advanced computational methods. To tackle this problem, Skinnider and his team developed DeepMet, a chemical language model trained on the structures of all known human metabolites. This innovative approach treats chemical structures as text strings, allowing neural networks to learn biosynthetic logic similar to how they process natural language, thereby predicting the existence of previously unknown metabolites with remarkable accuracy. The effectiveness of the DeepMet model was validated through a rigorous testing process that mirrored their earlier success in forensic chemistry with the DarkNPS model for detecting designer drugs. In one notable experiment, the team waited six months after training their model to see if it could predict emerging illicit drugs; remarkably, the model successfully anticipated over 90% of new substances discovered during that period. Similarly, DeepMet generated structures for novel metabolites that were subsequently synthesized and confirmed in human urine samples and mouse tissue data. This capability has already led to the discovery of more than 40 previously unrecognized mammalian metabolites, demonstrating how computational predictions can guide experimental efforts to map the complete landscape of mammalian metabolism. Looking toward the future, Skinnider emphasizes that while chemical language models have successfully narrowed the vast search space of possible molecules, further improvements are needed to directly link generated structures with mass spectrometry data. His lab is currently exploring strategies such as self-supervised learning and utilizing unlabeled data to enhance these connections. He concludes by reiterating that the inability to fully measure small molecules in biological systems represents a unique opportunity to address fundamental scientific challenges through computation, a mission that continues to drive his research group forward.
Read the full video transcript
Uh it's my pleasure to introduce uh the next winner of the early career award. uh this one is typically given to the young researchers in the bionformatics or computational biology who completed their PhD uh in the last six years and this year winner uh is Michael Skinned and he was selected because of the great his great achievements in the very different fields of binformatics and also by his uh support or maybe contributions to the binformatics community. Uh a few words about his career. So he actually started with the single cell uh RNA seek data. He did nice benchmarking of the methods for differential expression analysis. There he also came up with the machine learning approach agur uh to identify um specific cells which are critical for the healing process after uh um after walking issues. uh he also uh then moved to uh masspec data uh to protoomics and uh later to the metabolomics and there he developed a a new let's call it metabolic language model deep math to dive into the deep to the dark matter of the metabolics data and identify uh the new or until now not known metabolites in the in the spectra. for the community. He also developed several art packages and um from the other achievements uh he also helped to defend and maybe rescue the MD PhD program in Canada. And as a last uh kind of piece of information uh um he was offered a faculty position at the Princeton University directly after a PhD. So, please welcome to the stage uh Michael Skinn Skinner. >> Thank you very much for that uh very kind introduction. Um I want to start just by saying, you know, how grateful I am to be to be speaking to you today. um this is really an enormous honor and it's made more meaningful for me by the fact that many of the previous recipients um of this award are people who I've really looked up to and whose work has been very uh meaningful and inspiring to me. So um I I thought I would tell you um a little bit about um our work on developing computational methods to help discover previously unknown uh mamalian metabolites. Um before I jump right into that though uh I thought I'd just talk a little bit about um you know my own background and how we came to be uh interested in working on this particular problem. Um I am of course a computational biologist by training. Um that training really encompassed many different kinds of biological and biochemical data from genomics to transcrytoics to proteomics in addition to metabolomics. So uh for instance I got my start in research as an undergraduate student um at Hamilton or a university in Hamilton Canada uh where I worked with uh Nathan McGarvey and my work there focused on developing computational tools to discover bacterial secondary metabolites or natural products. Uh these methods used as input genomic and metabolomic data sets. And so this is where I really first developed a deep interest in metabolism and not just metabolism but in technologies for measuring metabolism. Um notwithstanding that interest uh when I moved to the west coast of Canada to do my MDPhD I also moved into a different field. Uh so my work with Leonard Foster focused on mapping protein interaction networks using a technique called co-fractionation mass spectrometry. uh and this work culminated in the first uh tissue specific protein interaction networks in mammals. Um during that time I also spent quite a bit of time uh at EPFL where I worked in Gregar Cortine's lab uh developing computational tools to identify subpopuls of neurons that uh produce a particular behavior. this time using single cell and spatial transcrytoics data as input. And ultimately these methods pointed us towards a subt type of spinal cord neurons uh which can be uh activated through electrical stimulation of the spinal cord in order to restore walking in human patients who were previously paralyzed. Um at Princeton as I said my lab has really come full circle to to my undergraduate interest in metabolism. Um but now instead of uh focusing on bacterial secondary metabolism uh we're really much more interested in the mamalian metabolom and particularly in discovering metabolites associated with the development progression or treatment of cancer. So I tell you all this in part to explain what it is about uh metabolism that I just could not stay away from. um you know I I told you about very different um themes and ideas in genomics, transcrytoics and proteomics and these are of course three very different fields but um I would argue that they all have one important thing in common which is that in these fields it's become routine to measure essentially the complete complement of DNA, RNA or proteins in any given biological sample. And the situation in metabolomics is really very different. uh it is still today essentially impossible to comprehensively identify the small molecules in that same sample. So this seems like a very profound gap in our ability to study and understand living systems. Um it's a gap that I think is particularly exciting uh from a computational perspective because I would argue that this gap is not for lack of an appropriate uh experimental technique. Uh in fact with mass spectrometry we already have an analytical technique um that can detect or acquire data from thousands of metabolites in any given biological sample. And not only that but massspec uh collects very rich and multifaceted data for every one of these uh small molecule metabolites. So this seems much more like a computational problem. How can we take this rich mass spectrometry data and deduce from it the structure of the metabolite that was observed by the mass spectrometer? And that unfortunately turns out to be uh sufficiently difficult that today uh in a uh routine metabolomic study really only a very small fraction of the signals that are required by the mass spectrometer will actually be identified that is to say linked to the chemical structure of the corresponding metabolite. And so the rest of that data is is commonly referred to as the chemical dark matter of the metabolome. So this uh dark matter is really what we're uh focused on illuminating. Um we're broadly interested in using uh computational tools, many of them leveraging advances in AI um to um discover this hidden chemistry that's sort of hiding in our routinely collected metabolomics data and really uh do that as a means to finally uh put together a complete map of the mamalian metabolism. So that is quite a grand uh challenge and you know you might ask how can we possibly make progress uh on this. Um I'll tell you a little bit about um the logical starting point for an approach that we've had a lot of success with over the past uh 2 years or so. And uh that that thought is sort of as follows. Um over the last century uh scientists have amassed an incredibly rich understanding of the known mamalian metabolom. So can we now leverage that understanding in a systematic way to predict the composition of the as of yet unknown metabolom uh and the computational approach that we've been using to explore this idea is based on uh chemical language models. So the core uh concept here is that we represent the chemical structures of known metabolites as short strings of text. We particularly like a format called smiles that I'm showing you here. And of course, we do this because representing metabolites as text allows us uh to repurpose the same kinds of neural networks that have been so successful at learning the syntax and semantics of natural language. Uh and instead of applying these language models to the words in a sentence, we can instead effectively apply them to model the atoms and bonds in a chemical structure. And so our hope was that this approach might allow us to as I said learn from the structures of known metabolites uh to generate structures uh for as of yet undiscovered metabolites which we could then target for discovery experimentally. Uh I should mention that this was an idea that we already had a certain amount of proof of concept for uh in a related field specifically that of forensic chemistry uh where we had previously used language models to uh helped discover emerging uh elicit drugs of abuse also known as designer drugs. So a few years prior to this uh work on the mamalian metabolom we had trained a chemical language model on the structures of all known designer drugs a model that we called dark nps and after we trained this model um we did something a little bit unusual in the field of machine learning which is that we waited um we actually waited for 6 months and over that period of course um new drugs of abuse were emerging on the market around the world and they were being discovered by forensic laboratories. And so at the end of that six-month period, we asked how many of these emerging drugs of abuse did our language model successfully anticipate. And remarkably, we found that uh dark NPS successfully generated more than 90% of these emerging elicit drugs in this perspective test uh including the uh structures that I'm showing you here. We then also showed that we could integrate the language models predictions with mass spectrometry data and that this allowed us to discover previously unknown elicit drugs uh in clinical samples as well as in law enforcement seizures um with really a remarkable degree of accuracy. And in fact, we worked with the Danish National Forensic Lab uh to discover or help discover a new uh dissociative uh drug, this derivative of uh the well-known street drug PCP that I'm showing you here. So um all of this you know past work really led to a lot of optimism that we could use essentially the same approach and rather than learning the medicinal chemistry logic of illicit drug synthesis we could instead learn the biosynthetic logic of metabolism. So uh we implemented that idea in a chemical language model that we named deep. uh deepmet is a model that's been trained on the structures of all known human metabolites and which as I'll show you can guide the discovery of novel mamalian metabolites uh and much of the work that I'm going to show here uh was led by Tony a PhD student in my lab. So I mentioned that um we had trained deep so that we could generate new metabolite like chemical structures which we could then target for discovery and um we found that deepmet did an excellent job of generating metabolite like structures. Uh so good in fact that a second machine learning model could not tell apart the generated structures from those of known metabolites. And just as we had seen with dark NPS, um we found that deepmet successfully generated the vast majority of newly discovered metabolites that were added to metabolic databases after we had trained our model. Um these predictions in fact were good enough that we could buy or synthesize chemical standards uh for compounds that Demet predicted ought to exist. Uh and then we could use the data that we had acquired from these standards to actually discover these metabolites uh in human samples. So here I'm showing you two examples of metabolites that DeepMet predicted likely existed as human metabolites which we were able to experimentally confirm in human urine. Um we also again developed approaches to integrate these predictions more directly with metabolomics data so that we could uh assign structures to many of the unidentified uh peaks or signals uh within the mouse metabolism. So for instance here I'm showing you a uh mouse metabolite that's very specific to the kidney and pancreas. Uh we synthesized the predictive structure shown at the far left. Uh and uh we found that the experimental data from this chemical standard uh matched almost perfectly to the corresponding signal in the mouse kidney. Um this is an approach that we've now um been able to uh scale up to some extent uh to the point that we've now used deep to discover more than 40 previously unrecognized mamalian metabolites. And this is really exciting to me because uh it suggests that um you know we could really uh use this technology to accelerate the mapping of the mamalian metabolism. So you know why does this approach work so well? Um, I think a good explanation is that chemical language models effectively allow us to narrow the search space. Um, when we're hunting for a particular category of small molecule, chemical space as a whole is sort of incomprehensibly vast, but metabolite like chemical space or designer drug like chemical space are small subsets. Uh, and so by narrowing our search to consider only metabolite like structures, we've essentially made this uh problem a lot easier. That being said, I I think that there are still really important opportunities to do a better job of linking these generated structures to the mass spectrometry data itself. And so this is really uh a major focus of the lab going forward. We have a few different ideas about how to do this that are now underway. Uh for instance, collecting more data ourselves, learning from unlabelled data, for instance, with self-supervised learning. uh learning from new types of data that are not really typically considered in structure annotation in metabolomics and then of course uh simply developing better machine learning models. So I'll stop there um maybe just by reiterating this idea that um our inability to comprehensively measure small molecules in biological systems is I think um an opportunity to address a fundamental challenge using computation and this is really the the challenge that uh motivates uh my lab. So um I'll I'll stop just by thanking um the people in my lab who worked on this of course our funding sources and once again I really want to express my gratitude to the Swiss Institute of Bioinformatics. Uh so thank you. >> Thank you Michael. Congratulations. Uh any question the audience? Yes please. >> Yeah thank you very much for the interesting presentation. Um I like the test that you did you know just waiting for design of drugs to be created and see whether or not you could have detected them and 90% is quite impressive but then my question automatically is actually how many of them did you generated and linked to that um since those are L&Ms I guess you are able to rank which are the most likely also in term of metabolites >> yes that those are all that's a a great series of questions. Um I sort of you know uh simplified for the purpose of uh just time but in fact we do um you know assign uh a score to each of the generated structures that sort of correlates well with its uh likelihood of uh occurring in the future. So the 90% number uh you know it reflects uh the proportion of generative structures that were detected but the denominator here is huge. It's uh on the order of millions. Um that being said in in both that work and then in the deep work we found that we could actually use that score to sort of prioritize without any analytical data at all which of the generated structures are most likely uh to emerge or be discovered in the future. And so I I imagine that's how you're trying to detect the new metabolites just looking at the most likely ones. >> Exactly. And then using that same uh idea to sort of uh reank our interpretation of the mass spectrometry data particularly the MSMS data itself. Yeah. >> Um sorry our schedule is quite packed. We don't have time for another question. I think that you could discuss during the lunch. Thank you. Um thank you again Michael.