Submind YouTube summaries
Thumbnail for nf-core/animal-genomics: Miguel Pérez-Enciso

nf-core/animal-genomics: Miguel Pérez-Enciso

Watch on YouTube

Video summary

Miguel Pérez-Enciso from CRAG delivered a seminar exploring the impact of artificial intelligence on animal genomics and breeding, focusing on genomic prediction, interpretability, and generative AI. Regarding genomic prediction, he observed that while AI has influenced phenotyping and aquaculture monitoring, early comparisons between standard linear models like GBLUP and deep learning approaches such as CNNs and MLPs have not yielded dramatic improvements in predictive accuracy. He argued that the current success of large language models relies more on brute force computational power and massive parameter counts than on fundamentally new algorithms, suggesting that future progress will come from integrating classical models with diverse data sources like climate information and satellite imagery rather than replacing existing methods. The discussion on interpretability highlighted that deep learning remains a black box where attention scores in transformers can only vaguely relate to genomic relationships or GWAS P-values, becoming increasingly uninterpretable as model complexity grows. Pérez-Enciso warned against relying on parameter choices in visualization tools like UMAP that allow users to shape data according to their expectations, noting that this practice risks biasing scientific conclusions. On the topic of generative AI, he described it as a simulation tool capable of creating new proteins, images, or phenotypes based on conditional inputs, such as generating fruit shapes from DNA sequences. He confirmed that while generating multi-omics data or whole genomes is theoretically feasible, it requires substantial training data to be effective. Pérez-Enciso concluded that AI will likely coexist with classical quantitative genetics rather than replace it, drawing a parallel to how books and e-books share the same space. However, he expressed significant concerns regarding the reproducibility of deep learning models due to their inherent indeterminacy and the risk of researchers becoming overly reliant on AI for coding and writing, which could compromise human learning and scientific rigor. He emphasized the critical need to maintain a balance between practical experience, peer validation, and new AI-driven knowledge to ensure the integrity of scientific inquiry in this evolving field. Following his presentation, the session transitioned into a seminar by the Farm Variant Function Task Force focusing on methods, beginning at 5:00 p.m. European time. The first part of this subsequent session features Fin Gray from the University of Edinburgh discussing high-throughput phenotypic screens in livestock species, followed by Shiva Abdess from Google DeepMind who will cover advances in regulatory variant effect. These topics continue the broader exploration of how advanced computational methods and data analysis techniques are being applied to solve complex challenges in agricultural genetics and functional genomics.
Read the full video transcript
seminar. It's great to see you joining us today. Uh and as always the goal of the goal of this series is to bring together researchers interested in animal genomics, bioinformatics, and related computational approaches, while also exploring new develop- developments and emerging topics across genomic more broadly. Before we begin, just a brief reminder for you all of you all for all of you uh that the session is being recorded and will be made available on YouTube uh for those who could not attend live. So, if you speak or open your uh webcam, you will appear in the registration. Today, we are delighted to have a new speaker with us, and I'm very much looking forward to the discussion that will follow the presentation. With that, I will hand over to Jose Espinosa Carrasco, who will introduce today's speaker. Jose, the floor is yours. >> You're you're muted. >> Yes, you're right. [laughter] The messages appear on my screen as well. So, yes. Uh hi everyone. Uh for today's meeting, I'm very happy to introduce uh the speaker Miguel Perez and Thiso from the Center for Research in Agriculture Genomics uh CRAG at the UAB campus in in Barcelona. Uh Miguel is a biologist who obtained his PhD in 19 uh 90 in genetics at the Universidad Complutense in in Madrid. And after his PhD, he moved to the USA USA and France for 10 years of postdoctoral studies, and he he specialized there in Bayesian statistics applied to animal breeding and and quantitative genetics. He then worked at at the Institute de Recerca i Tecnologia Agroalimentaria, IRTA, also in in Barcelona or in Catalonia at least from 1993 to 1999 and at INRIA in in Toulouse, France from 1990 and to 2023. When he became a at the moment he became a nuclear research professor and he's currently based at the center for research in agricultural genomics at the Huawei campus. Uh and between 2022 and 2024, he also was working for industry as a quantitative geneticist geneticist sorry for Corteva AgriScience. So he has also experience in in industry. And and he's now visiting the the CEREGE and there is where I I met him and he's actually a very nice guy. So Miguel, I'm looking forward for your presentation and the floor is yours. >> Okay. Well, thank you Jose. Thank you Francesca for the introduction and for inviting me to give this this talk. Um when Jose invited me, I I didn't have any really new new stuff to talk about, but I thought that maybe some thoughts on what is going on with AI and its impact that impact that is going to have and is having already in the industry and both in academia could be interesting as for discussion. So this is what I what I I This is what I I thought it was worthwhile talking a little bit or discussing with you. And I have arranged the talk in these few topics. So the first thing that really puzzles everybody who has a bit of experience on the on this area is why all all this fuss today. Why are are we everybody politicians, the people in the market everywhere everybody's talking is afraid of AI. Why is that different now? Because if you think it's it's has been just gradually what what is going on and I would explain that a little bit. But then focusing more on our topic of interest on breeding on genomics and so on I would discuss briefly the impact of AI on genomic prediction on interpretability also it's a topic that is highly highly discussed and highly disputed. How can AI be interpreted? Can Can it be interpreted? Does it make any sense to interpret these black boxes or these very very deep black boxes? And then I will follow with discussing a few issues on generative AI which has been perhaps one of the most striking and most influential AI topics all over the the place but also in in animal and and plant breeding. And to finish I would like to to give an example or whether AI will be replacing classical tools breeding linear models quantitative genetics even statistics. I mean how would be a statistics in a few years? Already mathematician has been deeply influenced by AI. Many mathematicians are using AI to prove theorems to to discuss as if it were a a very knowledgeable colleague. So what what can happen? Of course we don't know but at least we can dream or think a little bit what what can how can be can all this be in a few years from now? So just before starting just a few features of breeding animal and plant breeding. So as you know it has a very or a strong industry component. So this means that breeding the main data related to breeding is normally in the hands [clears throat] of private hands, not in academia. So, many of the interesting data are owned by by companies. And it's a very long-term game business. It's not that breeding companies they expect to have profits or a lot of profit immediately. This is like a very continuous very continuous activity. At least breeding is based on linear model theory. Nothing has changed there. But AI is is having an is having a lot of influence in especially phenotyping. Think of satellite imaging and also some new animal industries such as aquaculture, where everything is monitored and analyzed with very complex algorithms in order to provide the animals with the the most profitable environment and and genetics. And also we should be aware that we are under public scrutiny the more and more because of environmental concerns, also welfare concerns. But traditionally it was only because of the impact potential impact that breeding was having. But now it's also the impact of AI itself. All the huge demand on electricity, water for the big this big data data centers and so on. So, there is a lot of a lot of corners or issues that we need to to consider here. So, so why all the all the fuss? Why are we talking now about AI? Because AI is is is pretty old. I mean, already in the 1950s even deep learning resurrection, the the back propagation algorithm was already published in 1986 by Geoff Hinton who was Nobel Prize a couple of years ago. But why? Why Why is that? So so my my interpretation Okay, so my interpretation is that new large language models are actually uh generating language and text as if we could talk this with them as we if we could they could understand me and I could understand them. So they're making very very uh sensible uh language and text output. And we have always been told that if there is something that really separates us from the rest of uh animal kingdom is actually the language. The language is really uh a unique language property or so we thought. But the fact that then this is no longer the case, this is disturbing. So this is of course something that I think is making appealing the people or or disturbing the people or making people that this time is is really different. Of course, it's also different because now we can manipulate and generate any kind of data, protein structure, DNA sequence, music, text, etc. I I and I would like to argue that mostly most of the advantage is is is because of huge brute force, not because now we have very very different algorithms. I will I will show you just in a in a second. So for instance, in this slide you could you can see on the left I I think you can see my my pointer. Sorry, on the right you can have the efficiency of new GPUs, which are increasing dramatically more and more teraflops per per $1,000 for instance. But also another interesting thing is that AI has become a big industry. So, if a GPU processor maybe 10 years ago would cost a few thousand to thousand dollars, now they can cost up to maybe near $50,000 or so. So, this is not something that a normal academic group can buy. It is something that is more accessible only to to the big to the big industry, sorry. But now, what has happened also it what has happened is that there's a big inflation in the cost of all these um hardware, especially for solid disk, for for a storage. So, it's no longer than GPU um computer cost is going lower in cost all the time. It It is actually going up, and this is something that maybe you have realized already when you were wanted to buy or update your your your clusters on for instance. And here on the bottom side on the right, you can see the number of parameters that this parameter that these models are having, and this is really huge. Really, really impressive. So, here and then well, this is for the poor video gamers. Now, they need to pay much more for what they used to do a few years ago, but anyways. Indirect cost. Anyways, this is in this graph >> [snorts] >> you can see the number of parameters that most that many modern large language models have. So, for instance, for Claude, you they have up to 5 billion 5,000 thousand thousand thousand parameters. So, 5 to the to the 12th parameter. So, what is this? Well, it's a big number, but how big? So, just consider that Wikipedia contains about say five five five billion words and about seven million articles. And Wikipedia, we can think that it contains basically all all knowledge, most a big a very large part of the knowledge that now now we have on basically any topic, even biographies of of people, history, science, everything. And this is about 1,000 parameters per Wikipedia word, but about 1 million parameters per article. So, think of the standard methods like blob or regression. I mean, they have one, two, three parameters and with that we were able to construct very I mean, able to predict and to construct quite reasonable models. So, now we are talking on a scale that is 1 million 1 million larger than that. So, basically, what these models are doing are memorizing memorizing the whole knowledge that is around and then digesting it and giving a prompt that is translated into a numbers writing a reasonable prompt. But, I don't know if you know uh this absurd theater called the the lesson from Ionesco a Romanian French theater sorry, writer. So, in this novel in this theater, what happens is that the the the the the teacher is teaching a person who doesn't know anything. She she doesn't understand anything. So, what she does is to learn by heart all possible multiplications in the world. So, she asked him the professor asked, how much how much is uh 3,330 by 437.60?" And she responds, but she responds because she has learned by heart everything. So, large language models to an extent, they are a little bit like that. I mean, they have learned they have copied the whole knowledge that they have been trained on for instance about 4 months. And then they are able to kind of replicate, to massage the knowledge and provide a new a rewriting of what is already there. So, this is something that we should bear bear in mind. And of course, you are more or less aware of deep learning. Here, what I would like to um to insist is that DL algorithms are not plug and play. So, what does this mean? This means that uh you need some is not they are not easily reproducible because they are not like uh in in general accomplished packages. They are just something that are is staying there and on the web and is very often uh that the libraries are outdated. So, you can no longer run what has been proposed just a few years few years ago. And because this system is indeterminate, this has another um com consequence. So, the main consequence is that um in terms of interpretation, if we have many different set of parameters that can produce the same answer, so this means that they are really impossible to interpret because we have different interpretations. Every set of parameters values of parameters would result in different interpretations. So, this is also something we should bear bear in mind. And why uh how do these algorithms work? So, what computers of course only only understand numbers. I mean, we know very well how images are converted into numbers and and vice versa doing the intensity the pixel intensity is a number between 0 and 255. But how text are converted into numbers? So, this is something that is at least to me is relatively relatively new. How How can the Wikipedia be translated into arrays of numbers and how can be retrieved and so on. So, the the basic idea again is very simple. So, basically what the algorithms do is they separate the text into words into discrete units. This is called tokenization and then they are embedding embed. So, what is embedding? Embedding is an algorithm or is a many different algorithms that basically they transform a vector a sorry a text into a vector of into an array of n dimension whatever. And how is this done? So, basically this is done by looking at the words or sentences that co-occur across massive data sets. So, across all Wikipedia how many times linear or the linear word occurs together with regression word and so on. And again across massive data sets with massive number of parameters. So, at the end this is how these algorithms work. Basically by correlations so to speak. The interesting interesting aspect of this encoding is that they can be interpreted semantically. So, for instance, it's very famous uh the anecdote that if you have in a in an N-dimensional space and you plot the coordinates of So, for instance, Madrid, Spain, France, and Paris, they are close in the in this N-dimensional space. The word Madrid is close to the word Spain minus France plus Paris or other way around. Spain minus Madrid would have similar coordinates to the word France minus Paris. So, what I mean is that these coordinates have a meaningful semantic interpretation. So, when you are making a question to ChatGPT or whatever, this question is translated into a vector, into numbers. The The model looks for the most similar uh responses to in terms of these coordinates, numeric coordinates, and transform the obtained coordinates in back into text. Of course, it's very relatively simple to say, but of course, there are many many different >> [snorts and clears throat] >> say nuances, there are many difficulties, but what I want to stress is that it is mostly brute force, brute force followed by some intelligent algorithms, of course. Anyways, let's go or let's start a bit more concisely working on um breeding. As you can imagine, the most influence the most important aspect of AI initially was to in try to improve standard linear models. So, many people, including myself and colleagues, were working with trying to improve genomic prediction, blup, gblup, and so on. And as you can imagine, there is a lot of uh people working on these areas. So, for instance, recently there have been already about 500 papers on these on these area. Earlier work by myself, by Montesinos Lopez all in C mid some other Chinese groups, we did not find any dramatic improvement on predictability of convolutional neural networks and so on. And more recent work, I must confess that is is quite difficult to interpret with me by me because the the the algorithms have become so complicated that they're very even difficult to follow. The papers and the algorithms and they're also very difficult to to replicate. It is not a straightforward, but I suspect I always suspect when there is a some new algorithm that is really making it much much better than a standard linear linear models. So, my my feeling there, I am a bit agnostic. I am not sure that AI algorithms would really improve upon genomic prediction a standard genomic prediction algorithms like like G block or and so on. Well, a bit of uh uh commercial just in case you want to take a look at that. And in the in the one of the first works that we we did on the topic on the UK Biobank, this is some of the most important or summary results. So, you can see here different colors corresponds to different methods. So, for instance, black and grey are linear models. This is prediction accuracy, okay? With different models and different, uh, scenarios. So, black black and grey are linear models. Grey are kind of, MLPs, multi-layer perceptron algorithms, the oldest algorithms. And pink and magenta are convolutional neural networks. We did not test transformers at the time because they were not been invented yet, but to me, the main message is that there are no very big uh, differences in terms of predictive accuracy. And I am, as I was commenting, I am not really convinced that this is this is still so today. Nevertheless, I don't think that maybe algorithms need to really improve upon what is, uh, standard models [clears throat] now. They do not need to improve. What I think would be the future would be to combine these, uh, algorithms with, uh, new data information, like climate information, uh, satellite images, uh, growth, longitudinal data, and so on. So, maybe the idea is not to to improve upon GBLUP, but rather to integrate GBLUP into a wider algorithm, like an agent or something like that. I really don't know what how can it be in the future, but rather to integrate all sources or many more sources of information. So, that that that can be a way to include, uh, AI into, into, into prediction. Okay, so this is as far as as it goes with prediction, genomic prediction. How about interpretability? So, if you have been reading on AI, it has been only highly debatable, uh, topic. I mean, even with BLUP, deep blood, it's always been considered a a black box. So, imagine now these deep learning algorithms that are even blacker blacker and blacker. So, interpretability itself is a concept that is has not a single definition. It has many different metrics and so on. So, one possibility is to say, "Okay, so what? I mean, I just don't care what the algorithm is doing. I just make a question to chat GPT. It responds. It gives me a reasonable response. So, I don't care whether the algorithm is intelligent or not. I just wants a response." So, the same thing could go for for genomic prediction or a strategy for a breeding company, etc. I don't know how the algorithms would make use of the information that I'm I am provided providing, but it seems that the algorithm provides a sensible answer and it is helping me to to reach to reach a conclusion to act or to react to that. But, of course, we can think also, well, if we don't know what the algorithm is doing, if we don't know what are the main reasons why why are we getting this response to this question, then we will maybe we will never be able to go over over the the state of the art. I mean, we will never be able to to improve our knowledge to understand what we are doing or how can we what would be the impact of this mutation on what on what phenotype and so on. Why why would this happen? So, this is really a very interesting and and as I say, it doesn't have a single in my opinion single response. I would like simply to to show you some experiments that we were doing with transformers. So, in a genomic prediction uh scenario, uh transformers compute um the attention. And the attention scores is more or less can be interpreted of how much is important one is need in relation to the prediction provided by another snip. So, it is vaguely it is not exactly, but is vaguely related to the genomic relationship matrix. So, it's a square matrix that really tells me how redundant is this information, how helpful is information from one is snip given another snip, and so on. So, what we did is that um I run a small simulation to see where the attention is related to uh P value in a GWAS. Okay. And you will see that well, it is not a straightforward interpretation. So, here on this panel, you can see uh the P values of a given uh GWAS study simulated data where we have over 1,000 is snips, and I simulated a single QTL right here. So, GWAS really pick up the position, and there is some other uh noise here. So, this could be a false positive, and so on. This plot, the next plot here is what is the attention score for all these snips average over over layers, individuals, and so on for the same data set. So, it it is very interesting, in my opinion, because actually the attention of this is is really much much higher than the rest of uh is attentions um due to other is nibs. So, it would seem that attention is really actually could be interpreted as a something similar to P value. And you is is funny because even you have some fake attentions here coinciding with the fake uh false uh P values. And this is when we had a very very simple model. So, we have only one attention a head and one layer. So, a very simple transformer architecture. So, what happens if I add more heads, so more attentions within the single layer? So, for instance, here we have four heads, so a different model, a different transformer model. And suddenly you see that instead of being a a flat a flat line, it becomes much much uh shaggy, so to speak. And if I add more layers, then things become completely un uninterpretable. So, this is still uh coming out, but not very not very uh clear this. So, you can see that attention the the very definition of attention or interpretation of attention would depend exactly or is very very much on the model that we are using. And these models can have even very similar predictive accuracies. It's not that they're completely different in how they they behave. So, this is the same thing for the QTL. And with real data, this is with real data, a GWAS P value and with the transformer. Of course, we don't we don't know in with real data which are the true QTL. So, here the dashes are simply the [clears throat] peaks that are over 10 to the minus 10 to the minus three. And we see that the profile is very rough, but also that attention I mean, I wouldn't say that I could interpret this that this is Well, yes, this is a peak here, but not really not really a clear clear interpretation. So, my what I what I I guess what I think is that uh at best, transformer annotations can be interpreted only in very simple architectures. So, this is something that I think we should bear in mind when when we think about these architectures. So, the third topic that I would like to discuss very briefly is generative AI. So, generative AI is really like like the queen of AI. Everybody thinks or we think that is really an impressive tool and and I agree. And we all think that it's going to solve many of our problems. I think I am convinced actually that it has very many interesting and nice nice applications. So, for instance, a generative AI is is is a model is is like a simulation model. It's like a genetic simulation, but instead of having a set of parameters and a set of quantitative trait, etc. with given heritability, what really these models are doing are imitating. So, you given millions of photographs of famous people for instance, and the algorithm would generate new new photographs. And the algorithms are improving that much that you cannot distinguish often what is a which is an actual image from an image generated by computer. And you can do that with videos, with music, images, pictures, and so on. So, it's really having having a having an impact. So, there are many different algorithms. So, like for instance, generative adversarial networks, diffusion models are the most popular models uh these days. Autoregressive methods, flow-based, etc. So, there is a lot of literature on on the topic. I would say that the most important um topic or or or or variant of generative algorithms are conditional generative algorithms. So, algorithms that generate do not generate simply any any kind of distribution, but that conditional on on a given feature. So, for instance, we would be interested in [clears throat] generating images given DNA, for instance, images of fruits or and so on. So, not not simply replicating what we observe, but conditioning on um on a given on a given feature. So, for instance, at CRG, there are people working on how to generate uh new proteins that have improved uh features. So, for instance, improved uh characteristics that they bind better to a given ligand, and so on. So, not not simply replicating known proteins, but designing new proteins that have a given given uh characteristics of of of interest. So, these are called, as you can see then, uh conditional generative models. So, we applied some of these tools but in a very or relatively old algorithm to predict uh fruit shapes from DNA. So, this would be one example of application. So, we took This is using an old technology and I would now we would do it much more realistically uh with much more um accuracy, but simply to to show you what that what can you do with this this models. So, for instance, what we had is different picture from from tomato with their crosses and some SNP uh data sets um um characteristics. So, we trained the model on the features, the characteristic and SNPs and so on and the images one by one by the by the other. And then we predicted as as you can see here they are simply the profile. Now, we would do it with a complete picture, but uh 3 years ago or 4 years ago it was with uh I mean not not so perfect, but what you can see here is that the algorithms are able to predict This is removing. This is not used in the training, but when we give it to the training and this is what what what was predicted. And you can see that there is a close similarity between the shapes of the tomato predicted and and observed. So, in this one here, I don't know if you can see the the mouse, but this tomato strain has a lot of variety variability in the size. This is the core de boeuf, the ox heart variety [clears throat] and there are there were tomatoes of very different sizes. And what is interesting also is that you can go beyond DNA. You don't need to to look at DNA. So, you can I have images of crosses. So, for instance, this is one parent and another parent and this is the observed offspring. So, when we train the algorithm, of course, without using the the offspring or this particular crosses, we are able to predict the the offspring given these parent shapes. So, this means that basically shapes of tomato in these cases, they are they behave kind of additively. All right, there are also other kind of algorithms like variational autoencoders. So, in this case, what you what we you can see here is the different strawberries simulated with with a random well, kind of random shapes of strawberries. When we are using these kind of images as as training. So, the amount of possibilities of the number of possibilities is is very very large and actually sorry, generative algorithms are quite is a heavy area of of research. >> [snorts] >> So, as I was saying initially, can we think that AI will would supersede? I mean, will replace breeding, even statistics, quantitative genetics? Probably, you would say no. I mean, at least that would be my my my intuition. And this is what I still think. It it is going to be very very very difficult. The same way that books have not been replaced by electronic means, but rather they cohabit. They they they coincide. But nevertheless, in the last few years. Uh we have seen that quantitative genetics and breeding has become more and more computer-based and less and less on theory. So, if this trend continues, uh we can see that the role of AI will really uh trying will be much much higher than is than is now. And I will show you an example to see to see what do you think? So, uh we have seen the change already. I mean, we have seen that for instance, principal component analysis has been replaced by UMAP when using um single-cell data. Why has this happened in just a few years? So, in just the three years, people has started to do single-cell sequencing and instead of using PCA to principal component analysis to replace uh to visualize, sorry, the uh the data, people has started using UMAP. And UMAP is really uh well, also it's a visualization technique um proposed by Geoff Hinton among others. And yes, I agree that it produces very nice very nice plot and painting and so on, but what people don't know or or at least most people don't know is that UMAP depends on very many uh on different parameters and that the choice of parameters can have a huge impact on what you visualize. And whereas principal component analysis does not have does not depend on any on any parameter, well, only on how many eigen values and eigen eigen vectors you want to to plot. But uh UMAP depends on several parameters. So, the main ones are the number of neighbors, which represent how the method is going to attend or to focus on the local versus uh global structures. So, here you have the same data set represented with different values for the UMAP number of neighbors. So, you can see that these are the same data exactly as this one or this one. So, you can if we go back just a second, so you can simply realize that this plot that you you see here depending on the parameters that you are using, then you will find completely different um completely different uh portraits of of the data. So, what happens that people will choose the parameters according to what they expect to see. So, if these are So, for instance, if things are introns and these are exons and these are I don't know uh transposon, UTR, whatever, people would choose the parameters when they provide an arrangement that they like it. So, this is really really dangerous, I think, uh in this area. And another parameter is the minimum distance. So, again, this uh the the imp- the impact of these uh parameters is is going to be uh very large on what we observe. So, this is a method based on deep learning, basically, that we don't understand how it works, that has become extremely popular that everybody is using and that um none of us perhaps understand very well what it does. And none of us think of which would be the the parameters that that set that would be most more appropriate. Maybe there is none, of course. Maybe that can that can happen, of course. Okay. So, just to to give my my final thoughts on on on this topic. Well, as you can see, there are many many different aspects that could be discussed more more in depth. In my opinion, the the way I see it is that >> [clears throat] >> at least for the industry, there are many because they have the data and they have the power and and they have many people capable of of running these algorithms. There are many different opportunities to opportunities. So, for instance, one one application that I might think is developing different A I I agents, different algorithms, so to speak, start that they start to compete in each other as they do in general generative adversarial networks, for instance, to guide to guide breeding policies. So, they can they can think develop they can really they could develop algorithms that they they try to mimic what would be the best breeding strategy and see what would be the main one, but not deterministic. It's rather when you have different chat GPTs talking between between them and so on. Of course, another very important application are local large language models that they use only internal information, so that you can retrieve easily information and you can store information and retrieve it easily and in a mini mini meaningful meaningful way. There are There are also something called world models, which are models that basically try to replicate as realistically as possible a given a given a scenario, given a field. So, for instance, a field in a given or a series of fields or planting fields or they would be like like twin like twin models more or less. So, there are many different opportunities. In academia, I I mean in terms of research, I think that we are trying to understand what these models do. They're trying to interpret to improve perhaps some of the algorithms and so on, but it's still I think we are kind of drowned by so many so many new algorithms that appear every day. So, it's difficult really to keep up with uh with what has been what is being produced. Well, as I mentioned, precisely one of the most important areas, I don't think that AI or deep learning has produced any real improvement upon upon traditional breeding or predictive algorithms like GBLUP uh and so on. But, I think that deep learning would be very very useful if we can unify highly heterogeneous data and be able to produce to produce an improved or to or to or to to produce a synthetic a synthetic prediction that takes care of all all aspects of the environment and the genetic and the genetic information. Now, a little bit on the risk of the very big risk that I think are associated with AI, not especially for for inbreeding, but in more generally. One important thing is reproducibility. Science in general is based upon the fact that reproducibility is something given or is something that is important for the advancement of of knowledge. Now, I think this is at the stake because algorithms are so Someone is talking there or Okay. Or um I think this is something really really serious. Another important aspect that we have seen already is what would be the impact on human learning. This is something that I cannot it cannot escape from from my mind because we have all of us have become much more lazy for programming for instance. Even if we find the the slightest question, we just go to to cloud, to whatever language model and we ask, "Okay, give me a code to do this plot with these colors and with this data frame and so on." But my question or my my my concern is whether this will go beyond that. I mean, whether because as I as I was mentioning, these models what they do is basically they memorize the whole world knowledge will make us will kind of compromise our progress beyond what is stated. I don't know. As I mentioned, mathematicians are using these models to prove new theorems and to find even to find new theorems. So, this may not be the case, but there is there is an issue. And of course is having a huge impact on writing and reviewing. I don't know in reviewing how how the impact would be. I have never used it for writing reviews and and so on. I cannot think that this will replace this part of aspect. But of course as I writing well, we all know it has really dramatically improved in in this case. Science writing. So, my general conclusion would be it's it's a very nice time for action. I mean, if you want responses to a specific questions, they are there. I mean, really meaningful questions, responses and and so on. But for understanding how we go to how we obtain that question that sorry that that answer maybe is is so not so not so we we live in a pragmatical rather than theoretical say era. So, that would be more or less my my the way I see the way I see the impact of artificial intelligence modern artificial intelligence on our on our field, not only as you can see in breeding, but basically in any aspects of of science or even the the society. So, well, with that I I am done. If you have questions, you can do it through the chat or directly. If I Whatever you you prefer. >> Okay, thank you so much Miguel for this talk that for sure gave us some food for thought. Uh we have three questions in the chat and we also are running a little bit late on time because we have to finish finish at 5:00 sharp for the the seminar of the Fung Task Force. So, the first question is by David McHugh. So, David, if you want, you can unmute yourself. Or other otherwise, I will read it from the the chat. >> Oh, sorry. Thanks, Francesca. Can Can you hear me okay? >> Yeah, we we hear you. >> Can I Okay, yeah. Yeah, I was just wondering what Miguel thought about um you know, incorporating other omics data types, but obviously, it would be very useful to have those generated for the same animals that you would usually be using for genomic prediction and ultimately genomic selection, but that's not really feasible. But if you have data from smaller smaller numbers of animals in appropriate experimental contexts, um you know, gene analyses, functional genomics experiments, using the appropriate tissues that might relate, for example, to meat production, you know, muscle tissue or whatever. Um you know, can you see, you know, how will AI-based approaches, AI models, really help in terms of using that data in a meaningful way? What What does Miguel think about that? >> Yeah, okay. Yeah, I I I I think that using multi-omics data would be very similar or similar conceptually to what uh so for instance um climate data, environmental data, soil data, and so on. The The difference I see it as that many omics data, they are measured in a few animals or few individuals, whereas genomic data is available for everybody, not to mention phenotypic standard phenotypic data. So for that, I think is is is really something that is not only related to to AI AI. So how could AI I mean, to me the way is could AI impute the missing data? So that could be one one one question. I mean, that would be I think the the main question. So that I I I I I cannot I I cannot I cannot say. I think is is whether if the data are there, they are easy to to integrate. The way that I think this deep learning algorithms have not thought very much is about what happens with missing data because with missing data, normally you you throw away the whole row. If you don't have all information, they throw whole row. But it of course we need we need methods I think to to predict what is missing there. Yeah, so Mhm. >> O- Okay. Tha- Thanks, Miguel. >> Yeah, thank >> Okay, so the next question is is from Edgar Caballe- Caballero Vargas. So Edgar, I already see that you unmuted yourself. So >> Yeah, so thank you. Yeah, thank you. Thank you very much for this talk, Dr. Perez. So my question is basically like with conditional generative AI, I saw that you basically could generate the tomato likely phenotype. So I was wondering like related to the previous question like if you can integrate different like multi-omics data it could be like if you want to like know what it's this like like I want to generate a genotype, for example, that could have a potential. What type of like conditions could be could be like phenotype data or another type of I don't know multi-omics? Is that feasible? >> Yeah, yeah, that that's a very very very interesting project and actually sorry, question. And actually we we have a project we would like to explore these issues. There is some software, for instance, Evo3, I think, that is already able to generate complete DNA sequences but from prokaryotes individuals. But one possibility would be to use generative algorithms to generate the whole for instance, a SNP data and the whole associated um features. So, for instance, the the shapes of fruits, root architecture, or the conformation, color patterns from in dairy cows, for instance. Um yeah, that that I think would be feasible and and kind kind of doable. Um the question is to have the data also to be able to train the to train the algorithm. But it's something that in principle should be should be should be doable. Maybe not now, but maybe in in a few I don't know in in some time in in the future. Um yes, one thing that we were discussing in a paper with with Gustavo and Laura Fingaretti was how to combine also standard simulation with generative with the generative algorithms. And I think that they they can they can complement. Oh, oh, and you mean you mean also to generate or to combine all data. Well, in principle, you could you could simulate also omic data. That So, for instance, is very relatively simple to generate abundant data, microbial abundant data, in a few lines of code omic data, inspection data, and so on. Maybe is a bit more tricky, but uh Yeah. >> Okay, so the last question for today is from Andres Segarra. Andres, do you want to unmute yourself or should I read it? >> Hello, Miguel. >> Hey. Hi, Andres. >> I've been told that there's been a proliferation of grants with the use of AI. And the papers that I read that deal with AI, they are often full of cryptic models they don't understand and vague statements. And my impression is that essentially we are kind of back to 1960, where instead of reading papers, you just ask your friend or trust your experience. And this has to do with your last slides that you look you are very worried about the state the advancement of scientific knowledge. Yeah. >> Yeah, I I heard that So, for instance, they were an inflation of ERC grants and they were not able to review all of them. So, that's that's Yeah, that's really something that um concerns concerns me and and concerns us, I think. Uh whether what is the the what you see is is is is make sense or or is is is is credible or not. So, that's that's that's something that I still we don't know we don't have the the answer for that. But uh Uh, there was something who was saying, "No." These days you is not that I see It is not that I believe what I see, is that I see what I believe. So, when you have a pre-conditional thought or say an idea or could be even a grand proposal all of all of your effort is is going through through that. You only look at the data or at the uh, what supports your your your your your your opinion, so to speak. I don't know if you were talking about that, Andres, or um >> Yeah, more or less. >> More or less. >> [laughter] >> So, whether I trust my friend or my own experience. So, well, your own experience depends on your age. Um Yes, um I think that uh you will never have enough enough experience unless you you work in a very reduced area of of research. So, I think a balance you need a balance between uh your friends and and experience and and new knowledge from language models and so on. >> Okay, thank you so much again, Miguel. Uh, I hand over the the talk to Jose that we close it and thank you all for being here. >> Yes. >> All right. Thank you very much for the nice presentation and the beautiful discussion. Uh, I just want to announce that next seminar will be announced by the usual channels. Uh, and as always will be the third week of of July. We are working to confirm the speaker. And before we close, I also wanted to quickly mention that right after our session at 5:00 p.m. European time, there is an another seminar uh, that may be very interesting also for you. This is from the farm variant function task force that is kicking kicking off today and it's about methods. Today is the first session and and the speakers are Fin Gray from the University of Edinburgh who will discuss how to bring high throughput phenotypic screens to livestock species and Shiva Abdsec from Google DeepMind who style will cover how to advance regulatory variant effect.