Submind YouTube summaries
Thumbnail for MLFPM Symposium 2022: Joaquin Dopazo

MLFPM Symposium 2022: Joaquin Dopazo

Watch on YouTube

Video summary

Joaquín Dopazo, a founding member and thought leader in European bioinformatics, delivered an address focused on applying functional genomics and mechanistic modeling to the challenge of rare diseases. He highlighted that while classical research approaches often fail for these conditions due to their scattered nature and lack of market incentives for pharmaceutical companies, there is a need for a paradigm shift toward understanding shared disease mechanisms rather than treating each specific pathology in isolation. Dopazo explained that collectively, rare diseases affect six to eight percent of the population, yet only about 400 have efficient treatments available because scientists are too fragmented across thousands of distinct conditions. To address this, he proposed combining mechanistic models with machine learning to analyze massive datasets and identify common biological pathways that underlie various rare disorders. The core of his presentation involved demonstrating how computational models can simulate cellular functions by mapping biochemical pathways as electrical circuits where proteins interact to trigger specific cell fates like proliferation or death. Using gene expression data, these models allow researchers to predict the outcomes of genetic knockouts or drug interventions without conducting immediate physical experiments. Dopazo illustrated this with a study on kidney cancer and glioblastoma resistance, showing how simulations could identify why certain cells survive treatment by expressing alternative proteins that bypass inhibited pathways. This "revenge of bioinformaticians" approach allows for *in silico* hypothesis testing followed by experimental validation in the lab or clinical setting, effectively bridging the gap between theoretical models and biological reality before resources are spent on wet-lab trials. To further validate these predictions at scale, Dopazo presented work utilizing a massive digital health database from Andalusia, Spain, which contains detailed records for millions of patients. By analyzing this data within secure computing environments to protect patient privacy, his team successfully identified treatments that were effective against COVID-19 and other conditions by predicting how specific drug targets influence disease maps. The results showed strong statistical enrichment between their model predictions and actual clinical outcomes, proving the efficacy of using existing experimental data to formulate hypotheses without needing new experiments for every step of discovery. He concluded by outlining a systematic framework where partial disease maps are built from known mutations in pathways, allowing machine learning to propose lists of potential drug candidates that can be validated experimentally, thereby accelerating therapeutic development for rare diseases through shared mechanisms rather than isolated efforts.
Read the full video transcript
this is my great pleasure to introduce York Hindu puzzle your kin is a founding member even of the first of our two ITN so he he played a key role in these in these two two networks I mean it's almost unnecessary to introduce him he's a thought leader and a from a prominent person in in bioinformatics he's interested in functional genomic systems biology mechanistic modeling of or mixed data and their exploration he's the director of the FPS in in Syria and one of the key figures in bioinformatics in Europe we are very proud that he was part of our two networks and that he's opening our two days in posium here now with course and modeling and machine learning applied to massive truck repurposing in real diseases welcome thank you very much Kirsten so uh thank you very much to all of you for for having a for having give me this opportunity of being I mean opening this this uh last session so we were commenting yesterday night that it was a Pity that this cross in in this in this uh in this idea because um I mean especially For You students is networking is very important so we recommended yesterday that it's not I mean you have to be good because you have to be good for your future but but it's not only that people must know that you are good so that means that you have to do networking you have to be um you have to make you uh uh visible in this in this field right so you have to it was a Pity that you that you couldn't make the network in that these itns are designed too but I mean try to squeeze this last this last meeting that we are having here to try to do the last Network here and I remember that you have to make the network so um I'm gonna talk you a little bit about some of the thing that we are doing in the I mean I'm going to be to make a mixture of of things right so I'm gonna talk about rare diseases but I'm simplify some of the methodology that you that you use with uh kobays in which we have been able of demonstrating the efficiency of some of the methods that we are using so um do you know that we are part of the Spanish Network for rare diseases and we our contribution typically is related to the analysis of of uh uh and the diagnosed cases um typically this patient for which you have an exome or a genome and there is not any mutation which is characteristic of the disease and so there are and they are not so they send this this uh hopeless cases to us so we do some research and we have a 30 percent of uh case resolution which is quite okay uh combined to the literature but now this is um I'm going to present you something which is more philosophical it has to do with the way in which research is done in in rare diseases so these are some facts on rare diseases uh so by definition is considered rare analysis if it affects to less than one person among 2000 many of them affect even less people maybe some of ultra rare diseases affects uh one person among two million Etc so typically there are diseases which are thought to be um rare by definition but at the end there are seven thousand actually there are some literature that say that there are 9 000 rail diseases so collectively they affect to six to eight percent of the population so it's like I mean collectively taking collectively is like a prevalent disease a normal regular prevalent disease right so most of them have genetic bases and there are very few treatments available for for only only four 400 rare diseases have an efficient of more or less efficient treatment right so what are the consequences of this mainly that the classical approach that we use for facing diseases is not very useful in rare diseases because typically you know there are a lot of people working on Diabetes working on a lung cancer working or whatever but we have maybe at least seven thousand this is so we don't have a lot of sets we don't have 7 000 sets of scientists A group of scientists working in any of them so what happened is that at the end uh the research is very scattered uh obviously in terms of treatments is the same Pharma companies do not invest in rare diseases because if you consider one by one you consider a treatment for one specific disease thank you uh I mean there is there is not a niche of market for them because there are very few patients right so at the end we should change a little bit that way we do research and the way in which we try to cure rare diseases maybe what I'm proposing is not the Panacea but when we need a change in the Paradigm so um we need firstly probably to focus on disease mechanisms more than on specific diseases because many of these diseases actually they share some of them this is mechanisms what is the advantage the advantage is that in that way all this knowledge will that we gain on mechanisms will break the disease barrier the problem is that we don't have much detail on on mechanisms of running diseases because they are mostly unknown right so um well we try to use mechanistic models combine it with a console machine learning to try to to gain some inside in in this in rare diseases from the point of view of the translational application uh we are going to focus on drug report policing why because because in that way we will be dealing with with drugs for which the security profile action mechanisms action Etc is already known so the only thing we we have to do is to prove that it's efficient in this disease and a lot of steps in the regulatory parts of the of the of the approval for for a dry is already solved right um so at the end the problems is that the relationship between the targets of this uh already drugs already in use and the disease are unknown what solution can we use again um some Coastal machine learning in combination with with that with mechanistic models so I shown that in a previous talk but it's good to refresh because probably you don't remember uh what are mechanistic models so we we use mechanistic model to provide a quantitative representation of the uh functionalities of the cell basically what we do is to use uh pathways this biochemical Pathways which we have the relationship between the functional relationship between proteins how proteins interact but not only physical internet but functionally interacts to each other and what this product is does at the end of this of this pathway so in some cases they trigger uh cell deaths in other cases they didn't trigger I mean a lot of I mean the the way in which the cell uh decides what to do the fate of the cells right so something which is interesting is that in this pathway we can Define uh functional endpoints so at the end of all these Pathways there are a function which is triggering the cell and the idea is to have something some mathematical modeling that we can that we can use so that would be the the framework and how GS interact with each other and we can use we can use data measurement that we have on the on the condition of the cell and typically we use gene expression which is a data which is nowadays is quite accurate and it's cheap to obtain there are lots of this data and they are sort of read out of what the cell is doing in this moment it has like a snapshot of what the cell is doing in this in a particular condition so that would be a very simplified picture how one of these models works so this would be the pathway would you have an interaction between some proteins on the left that receives some input some signal or whatever and they communicate to each other like in a circuit and finally they trigger a function and there are other type of Pathways which are the the um metabolic pathways in which the functionality would be the generation of metabolite but conceptually are more or less the same so the idea is that we put uh data in expression data that we mentioned in different conditions on this map and we see what happened there is a Formula which has a recursive formula which by means always we can say okay what would happen if there is a signal here and we have these states of activation in that case it would the the bulb with light so it's at the end it's like an electrical circuit what we are simulating here and what would happen in the second condition the second condition we have a short circuit here so the the ball will not lie so this is essentially what we have right um if you take real Pathways and you take real data you get things like this so this is uh real data with the tcea of the cancer adenone project and we put the data on the on the modeled pathways and we see that for example some functionalities that you can easily map to a whole cancer Hallmarks for example DNA replication which is how the cancer I mean grows up so if you put you take sample from many biopsies of cancer in that case is kidney cancer I think we just hearing kidney cancer because there are lots of data of kidney cancer so you measure what is the activity that you infer from the model based on the gene expression and you do a plot of survival plots you can see that patient with a DNA replication highly activated they have a bad furnaces but significantly bad progresses so this is very significant but you you can identify other currency whole numbers like and the apoptosis so patient with the antiboptosis activated they die more patient with uh an inactivation or celebration which means um metastasis half a barcelonosis as well patient with activated angiogenesis die more so at the end if we identify uh in that case cancer Holman but in in other cases you identify some functions which are related to our disease or to our phenotype we can easily follow the activity State based on the gene or the gene levels the inactivity simply right but there is something I mean this is good but there is something that I like still more of these models these models are very interesting because you can simulate the conditions that doesn't exist yet so you can take a condition and you can simulate a new condition in which you knock out a gene for example and compare and say okay what is the difference because between the previous condition and the condition with the knockout let's see what happens so you can sort of forecast what would happen in different situations you can simulate for example the the activity of drugs or whatever you can read probably you try to simulate something that 20 Knockouts and 15 over expression probably you will this I mean distort the system too much and maybe it's not um I'm not sure how about the result but for just one or two Knockouts it could be interesting so at the end what you do is to take the first condition and to do Knockouts how do you do a Knockouts here so is very easy simply you substitute the value of gene expression by zero or by a very long amount in that way you simulate I mean a protein in activation is like is actually like a protein knockout like I've been removing removing the gene there so you can Force this situation in which uh you have no cow that doesn't affect the function that really do affects to the function so you can before doing the experiment you can sort of see what would happen and actually this is what we did and I like very much this this uh this paper and I always say that once I publish this paper I can't retire myself very happy because this is what we call the Revenge of the bioinformaticians this is one case in which we uh came up with the hypothesis we tested in siligo there are no accounts and then we went for a experimental group and said please can you test if this is real or not so they tested that it was real so we we uh sort of uh predict the uh the the no cows that will I mean not kill but clearly damage the the ability of the of the cell to to replicate so this is something that can work and actually we publish it in in cancer research so we went to the to the battlefield to the cancer Battlefield and we we we won with our tradition we were already accepted in the beginning but it was nice so I mean these are another very nice example in which what we did was to use single cell to see what happened in a population of cells and in that case is the um glioblaston because typically you give the treatment which is the baby system up typically you remove the cancer so there is no cancer visible and in nine ten months the cancer goes again that is a residiva and the cancer is resistant to the treatment so they say oh the cancer was uh has acquired this resistance so what we so here is that if we simulate The Knockouts uh caused by bebasi sumac in the in the cell population what we see is that most of the cell were killed but a few of them were not killed and why they were not killed because this I mean this baby system map is mainly uh anti-back this uh this protein this protein bath is in the beginning of the birth pathway what triggers a lot of processes related to cell proliferation so what happened in the resistance cells is that they don't express that they are not expressing that actually they are expressing this other protein this PDA gfd that also triggers the the same pathway so you are trying to uh inhibit A protein that is not there so probably this is the silicone and the silicone was there and was not very successful so what but once you kill the smart clone then the silicone promotes grows up and then when when it takes over the the brain again as you try to kill them with an anti-bath but they are not using back for for growing up so this is a very nice way in which you can you can dissect what is going on at the cell level so uh yeah but what's the problem with this approach the problem with this approach is that that we relay on uh this circuits that are drawn by the in the in the pathways so the problem is that only one third of the genome is part of these Pathways right so um two-thirds of the genome is if they are relevant for the disease we cannot include them in the models it's well I must say that this this very nice cartoon was made by sankot who was a our student from the previous uh ITN ipn Network and I use it very very much so the we have this problem we have this problem with we need a lot for this modeling you need a lot of information biological information we don't have for for members of the deal so um the problem is that the generation of this biological knowledge I mean the drawing one of these arrows is a lot of time for typically this generation of biological knowledge is love several Laboratories working and trying to demonstrate that this Gene actually is uh um doing an activation of this other Gene and then you can finally draw a line a council line so it's this Gene is calcium phosphorylation for example this protein or the other protein or whatever right so um this is a problem we cannot wait for 50 or 100 years for all these sort of these arrows are drawn so one thing that we saw this uh would it be possible to use my children in to learn biology to say okay let's let's let the system to learn biology yeah well the problem is that machine learning has been applied in many scenarios in which you have this very good balance between the variables that you have to learn in the sample in biology we still I mean in terms we try to learn only biology meaning all they can be all the possible interaction direct interaction between proteins we still have a lot of variables and we still have few samples comparatively and actually um it's not only a matter of course or dimensionality is that the relationship between the genes are much more complex like the relationship between pixels in a picture a picture in a picture are related with the pixel pixel around and makes sense with a pixel around but genes are have crazy connections sometimes so it's not it's not it's not a problem on the same level there's nothing that we can do is to reduce the missionaries of the problem so we are not interested in in winning all the careers for all the biologists in one month and discover everything but something that we could do is to say okay let's try to see if some protein which are of Interest or transition can be related to the current knowledge that we have so this is a problem which is uh affordable probably because because I mean the dimensionality is much slower so uh I'm gonna switch them to copies so we have some copied funds to do some things and something that we did was well we were participating in this um because this is mapping which they uh uh draw up a very nice very detailed map of the all the process of virus infection and all the consequences uh Downstream on inflammation on how the the virus trigger the immune system etc etc so in that case it was very easy because we have a very detailed map so we we will need to to do anything with the with the map of the disease so what we wanted to know is uh what are the connection between targets of other drugs drugs approved for other diseases what are the the connection between these targets and the disease map of copies and not to all the parts of the map but to a specific part of the map in which we were interested on so the advantage of having these maps and having these functions at the end of these Maps is that you can focus on specific uh parts of the disease and inflammation and immune system or whatever right so uh we model this um this this map and actually we have a one version of the of the modeling of this ipatia model that is includes specifically the the code map and what we did was to uh well we have Carlos here for details and you want to ask specific details on the on the methodology but at the end the idea was okay let's try this let's try to explain the what is the behavior of the of the disease map in that case of the Kobe's map as a function of the different targets of drives that are already in use we can manage to explain the the behavior as a combination of one or several several targets of tracks probably this this draft will have an effect on the map right that that was the idea so we use this shot uh Supply explanation to try to look for the specific relevances of specific variables in that case drug targets and well we draw some some maps of activity of this this we found different situations for example this is the famous chlorokine the famous chlorokine uh uh acts on the map but absolutely on all the map and many other parts so I mean it's like if you I mean uh probably Barney in a cell is very efficient but but it balanced to you as well so it's not I mean it's official for combating the disease but also to combat the the the the the the patient so so I mean we were focusing in a specific uh in drugs we were more specific of certain time processes and I mean we managed to to produce a list of of drugs and by this time uh we saw that in a publication that they did um on a review on and trials that were for for testing treatments and prevention of kawaii uh all the drugs that were in in trial for for kovid who has a known Target because there were drugs which were I mean there were trade them that were more specific like gas inhalation or whatever so in that case we don't know what is the target but for this drawers that adari was known uh all the drugs that were there were predicted by the method okay I mean that could be good but we wanted to to have a stronger proof so we use this um these database that we have in Andalusia is um is this I mean you know is it the south of Spain it's a large region in Spain and actually the third largest region in Europe it has a population of 8.5 million so it's I mean this is the same size of Switzerland or Austria I mean it's like a medium uh size country in Europe so and we have an advantage is that the all the health system is digitalized and all the health system dumps the data into a large database and this data [Music] um put in a way that can be I mean queries so there are structural data we have also unstructural data but we have lots of structured data so we have here 13 million people so probably if you're not the biggest is one of the biggest database with detailed clinical information we have so something that we did was to to look for patients here for coveted patients for the first way and I mean we started that in the from the first world so we have in the range of 17 000 patients and some of these patients have received a another treatment for other reasons the vitamin um whatever I mean all the treatments because they were having this treatment they they were infected and we compare what happened with this patient that were having this treatment we speak with patients that they were not having this treatment uh taking into account all the core variables right actually we managed to to make a very nice silk with because I mean you know that accessing clinical information is is not this in general is not easy because I mean it's protective uh I mean I mean obviously it has to be protected nobody wants to have their own life explosive I mean I understand that but at the same time it's it's a problem for for doing a lot of studies right so what we managed to do is hey what what is the problem the problem is extracting the data from the health system okay what if we don't distract the data from the health system so we managed to put some Computing facilities within the health system and we can then analyze the data within the health system we say oh that's that's okay that we are happy with this so we set up this circuit in which uh necessarily we propose um whatever I study we pass through the ethical committees we then write the the this sheet of evolution of impacted data protection and then they since there is no impact in data protection and have the approval of that the committee so they provide us with the data we can do the analysis and the only thing that we make public are the results right so that was very nice I'm not going to to talk about that but that it has been a a complete change in the way in which we can do research now in Andalusia and we are trying to open that to to everybody right so finally what we saw is that they were 21 treatments that were highly effective so they protect clearly to patients and actually there is one who is counterproductive right and actually for most of them we since we have also data on analytics so we can follow the for example the lymphocyte counts and we say we see how for this patient also the lymphocyt counts uh was compatible with uh I mean with an improvement of the health right so interesting thing is that we have an enrichment of a uh of the I mean I'm on this data we have a lot of prediction not all of them actually for example we we didn't predict the the first one was was not crazy but the second world was predicted so I mean that's for me this is this is the definitely proof that actually the predictions were I mean relatively good because most uh there is an enrichment here a statistics that of predictions that we made with using the model so if we then know that this model is good we are in a situation in which we I mean this is this is very nice because we have made all the roles from the of scientific discovery of the scientific method proposed by Galileo Galilei in which you have to formula the hypothesis do experiments and check if the experiments fit through the to the the observation of it to the to the hypothesis we can do that without doing any experiments why because there are lots of data available so we can do everything without doing a new experiment I mean it doesn't mean that that experiments are useless because this data were obtained by previous experiment but what I mean is that we have now so many data produced by experimental in many cases you have the data already there which is very important so just for finishing we apply this we are applying this concept to the to radio diseases and with the idea of of trying to say okay um instead of of [Music] focusing on diseases one at a time we are going to focus on disease as a particularity of the whole uh cellular mechanism and say okay these diseases are characterized by mutation these three genes these three genes are in this part of the pathway so we have a representation a small representation not complete but a small representation of the disease map of this disease maybe we have missing parts but at least we have a part of the disease and then we have uh it was a little bit more than 150 reduces which has mutation within the known Pathways so we managed to make this I mean like I mean trading all the diseases like uh a part of the of the of the cell Behavior so okay this disease is here this is here this is here and I think we have here I will show you another slide later so I'm running out of time or oh okay okay so um so the idea is then do something similar to this and let me show you them what I mentioned before using genus person data and trying to learn if any specific disease map can be explained by the combination of of uh Target of all that drives and the idea has to make this systematically so we have we have the the whole map and we map Disease by disease the genes in part of the map and we build up this small partial maps of the disease and we try to see what drugs could be acting there it's not perfect but it's something that can be done systematically and you can solve in one shot you can propose a lot of treatment for a lot of diseases so the idea is we we have the genes the general I mean the the current knowledge we have the specific disease map the models and we do the machine learning for any specific disease and we look for the most relevant targets that are affecting them and then we go for the validation uh interestingly you have a look at the oh sorry did you do this I mean this clusters you see that there are different subclusters so at the end with the decisions at the end um as we suspected many real diseases at the end they are they are sharing uh mechanisms so probably drugs are can be used for more than one rare disease so we have I mean a couple of validation so this have was published a couple of years ago or two treatments that we uh predict for uh Franconia anemia were validated experimentally and now they are a systematic validation of all the training that we proposed for already discrimentation we are working since we are working with the people in the in the Spanish Network for various diseases we are in collaboration with several groups that are doing the specific validation so it will take time but at the end what we provide them is with with some jazz instead of trying to see what would be the drug so there are there is a list of potential candidates that they can use to to start with uh well this is uh I'm finishing yeah this is a bit of publicity so we have some software that you can you can use it you want to use the models and this is uh the people on the our supporters and we just show you this last slide The Bicycle Workshop that uh well I mean this is the list of people some of uh so for example um Antonio is there so he's going to participate and some of you are attending uh this is another place which you can do networking which is important so thank you very much and you have a question I will be happy of taking them thank you thank you Akin are there questions from the audience Giovanni thank you for these very interesting talks I have a couple uh question on the whole talk um one of them is on these explanation methods like shop um I'm a bit familiar with the problem is that often these postdoc method are a little bit unreliable or vulnerable to other serial attacks or some other issues um how to say it have you found other than just making predictions and then validating the drugs have you tried to find alternatives to that like using multiple of these interpretation methods or something that could give you an idea before the experiments if um I'd say what you are hypothesizing as a reason to be considered valid yeah well I mean apart from Carlos can give you a later a more daily explanation from the point of view of the focus uh or why we focus first on on shop is uh I mean typically we go very fast so we need to solve problems quickly and that was um well I mean a simple way of trying to see what is the contribution of any of the variables that was more difficult to obtain from the from the from the model I I mean I don't see that adversary attacks here are really bad because it's not it's not the case here but it was it was um I mean simply um since we are doing a prediction based on predictions so probably we are not going to be very you know picky with with the methodology it was only uh the necessity of trying to figure out what of these uh variables was having a bigger or stronger effect on the on the pathways thank you and our curiosity if I can quickly before somebody else asks a question uh he should that he used for example cake Pathways and a few times I've looked into them and the thing is they are a bit of a hot part of jeans and metabolites and maybe sub Pathways how do you actually convert something so complicated to a relatively simple model like the ones you were showing yeah well I I didn't mention that so dealing with Padu is a nightmare because actually we are having problems of using Pathways because do not care now is has become uh not private but I mean you have to pay for advice or something this is a bit problematic so the point with with care is that um they they have they are um they are these metabolized but metabolize can be easily removed but typically they are uh meaning that they have essentially protests acting on other products we would have preferred to use for example react on because we have a lot of relationship with ebi so they are pressing us to use reactant they probably will react on either they have not only metabolize they have a lot many parts of the map are for example how a protein how different employees make a complex so all these arrows cannot be modeled because what we have is a is a snapshot of the genus pressure so the idea that we have is if we have all the all the proteins so the complexes the complexes there we need only one node with different proteins but so it's very difficult for us to convert all these um arrows that are not functional activities in the map but are other representation of the biological knowledge to convert because something like 50 percent of the arrows in a rectum are all this stuff and this stuff can I mean it's not useful from the point of view that we want to use the map that is to put a Genus person data and to see what happens right but this is an image you have to do a lot of it's not as immediate as putting the the map in the models you have to do a lot of manual creation thank you another question here okay thank you very much Joaquin for being here and for the great presentation and uh I'll keep it short in the interest of time but just two quick questions the first is going back to those mechanistic models that you showed at the beginning is there any work on longitudinal longitudinal aspects of these like how the connectivity evolves over time during development for example of organisms and the second is this huge database that you showed of healthcare data in Andalusia what's the prospect actually actually accessing that database thank you okay I'm gonna answer for the second one um we have a um uh some instruction for for using that this data so uh so it's something very similar to what I draw there so firstly you have to uh ask permission for the to the to the SS committee for most of the I mean if you if you're uh studies reasonable you will get the permission for sure and the second and most problematic step is to pass this evaluation of impact in data protection so typically you just feel a series of questions so is the data going out if it's going out how you um guarantee that the data is not spread out you you are not going to try to re-identify patients Etc so what happens is if you get out the data from the health system you check one of the one of the most horrifying checks and then you don't get the approval for for so what we did was to to set up that in a way which uh now is not perfect but it works so you have to ask essentially you have to ask us to do the job so what we are trying to do now is to habilitate a system by means of which uh once you get the approval of the headies committee you can access you can manage the data without having access to the data something like using a virtual a virtual monitor whatever we do you are you can do things but you cannot copy the file outside this is only a technical problem we are trying to see how to solve it as soon as it is solved probably it will be more open to you because we want to to I mean to become leaders in you know this in exploitation of of clinical data and the first question was I don't remember sorry was there something in the connectivity or what you know evolution of these models how the connectivity evolves during the development for example what connectivity that's represented in the mechanism the connectivity that's represented in the mechanistic models that you showed at the beginning I was wondering if there's any work on how those evolve uh over time during development for example um not as far as I know so I remember with either a very simple study but using enrichment uh genomatology so we saw how the function evolved a long time in a system like it was uh I think it's interesting to see how functions move across time but now I as far as I know there are probably they are some study but I don't know thank you very much Joachim thank you very much for asking the questions also um thank you for opening our Symposium um so a round of applause [Applause]