Video summary
Miguel Pérez-Enciso from CRAG delivered a seminar exploring the impact of artificial intelligence on animal genomics and breeding, focusing on genomic prediction, interpretability, and generative AI. Regarding genomic prediction, he observed that while AI has influenced phenotyping and aquaculture monitoring, early comparisons between standard linear models like GBLUP and deep learning approaches such as CNNs and MLPs have not yielded dramatic improvements in predictive accuracy. He argued that the current success of large language models relies more on brute force computational power and massive parameter counts than on fundamentally new algorithms, suggesting that future progress will come from integrating classical models with diverse data sources like climate information and satellite imagery rather than replacing existing methods.
The discussion on interpretability highlighted that deep learning remains a black box where attention scores in transformers can only vaguely relate to genomic relationships or GWAS P-values, becoming increasingly uninterpretable as model complexity grows. Pérez-Enciso warned against relying on parameter choices in visualization tools like UMAP that allow users to shape data according to their expectations, noting that this practice risks biasing scientific conclusions. On the topic of generative AI, he described it as a simulation tool capable of creating new proteins, images, or phenotypes based on conditional inputs, such as generating fruit shapes from DNA sequences. He confirmed that while generating multi-omics data or whole genomes is theoretically feasible, it requires substantial training data to be effective.
Pérez-Enciso concluded that AI will likely coexist with classical quantitative genetics rather than replace it, drawing a parallel to how books and e-books share the same space. However, he expressed significant concerns regarding the reproducibility of deep learning models due to their inherent indeterminacy and the risk of researchers becoming overly reliant on AI for coding and writing, which could compromise human learning and scientific rigor. He emphasized the critical need to maintain a balance between practical experience, peer validation, and new AI-driven knowledge to ensure the integrity of scientific inquiry in this evolving field.
Following his presentation, the session transitioned into a seminar by the Farm Variant Function Task Force focusing on methods, beginning at 5:00 p.m. European time. The first part of this subsequent session features Fin Gray from the University of Edinburgh discussing high-throughput phenotypic screens in livestock species, followed by Shiva Abdess from Google DeepMind who will cover advances in regulatory variant effect. These topics continue the broader exploration of how advanced computational methods and data analysis techniques are being applied to solve complex challenges in agricultural genetics and functional genomics.
Read the full video transcript
seminar.
It's great to see you joining us today.
Uh and as always the goal of the goal of
this series is to bring together
researchers interested in animal
genomics, bioinformatics, and related
computational approaches,
while also exploring new develop-
developments and emerging topics across
genomic more broadly.
Before we begin, just a brief reminder
for you all of you all for all of you uh
that the session is being recorded and
will be made available on YouTube
uh for those who could not attend live.
So, if you speak or open your uh webcam,
you will appear in the registration.
Today, we are delighted to have a new
speaker with us, and I'm very much
looking forward to the discussion that
will follow the presentation.
With that, I will hand over to Jose
Espinosa Carrasco, who will introduce
today's speaker. Jose, the floor is
yours.
>> You're you're muted.
>> Yes,
you're right. [laughter]
The messages appear on my screen as
well. So, yes. Uh hi everyone.
Uh for today's meeting, I'm very happy
to introduce
uh the speaker Miguel Perez and Thiso
from the Center for Research in
Agriculture Genomics uh CRAG at the UAB
campus in in Barcelona.
Uh Miguel is a biologist who obtained
his PhD in 19 uh 90 in genetics at the
Universidad Complutense in in Madrid.
And after his PhD, he moved to the USA
USA and France for 10 years of
postdoctoral studies, and he he
specialized there in Bayesian statistics
applied to animal breeding and and
quantitative genetics. He then worked at
at the Institute de Recerca i Tecnologia
Agroalimentaria, IRTA, also in in
Barcelona or in Catalonia at least from
1993 to 1999
and at INRIA in in Toulouse, France
from 1990 and to 2023.
When he became a
at the moment he became a nuclear
research professor and he's currently
based at the center for research in
agricultural genomics at the Huawei
campus.
Uh and between 2022 and 2024, he also
was working for industry as a
quantitative geneticist geneticist sorry
for Corteva AgriScience. So he has also
experience in in industry.
And and he's now visiting the the CEREGE
and there is where I I met him and he's
actually a very nice guy.
So Miguel, I'm looking forward for your
presentation and the floor is yours.
>> Okay. Well, thank you Jose. Thank you
Francesca for the introduction and for
inviting me to give this
this talk.
Um when Jose invited me, I
I didn't have any really new new stuff
to talk about, but I thought that maybe
some thoughts on what is going on with
AI and its impact that impact that is
going to have and is having already in
the industry and both in academia
could be interesting as for discussion.
So
this is what I what I I
This is what I I thought it was
worthwhile talking a little bit or
discussing with you.
And I
have arranged the talk in these
few
topics. So the first thing that really
puzzles everybody who has a bit of
experience on the on this area is why
all all this fuss today. Why are are we
everybody politicians,
the people in the market everywhere
everybody's talking is afraid of AI. Why
is that different now? Because if you
think it's it's has been just gradually
what what is going on and I would
explain that a little bit.
But then focusing more on our topic of
interest on breeding on genomics and so
on I would discuss briefly
the impact of AI on genomic prediction
on interpretability
also it's a topic that is highly highly
discussed and highly
disputed. How can AI be interpreted? Can
Can it be interpreted? Does it make any
sense to interpret these black boxes or
these very very deep black boxes?
And then I will follow with discussing a
few issues on generative AI which has
been perhaps one of the most striking
and most influential
AI
topics
all over the the place but also in in
animal and and plant breeding.
And to finish I would like to to give an
example or whether
AI will be replacing
classical tools
breeding linear models quantitative
genetics even statistics. I mean how
would be a statistics in a few years?
Already mathematician has been deeply
influenced by AI.
Many mathematicians are using AI to
prove theorems
to to discuss as if it were a
a very knowledgeable colleague.
So what what can happen? Of course we
don't know but at least we can
dream or think a little bit what
what can how can be can all this be in a
few years from now?
So just before starting just a few
features of breeding animal and plant
breeding.
So as you know it has a very or a strong
industry component. So this means that
breeding the main data
related to breeding is normally in the
hands [clears throat] of private hands,
not in academia. So, many of the
interesting data are owned by by
companies.
And it's a very long-term game business.
It's not that
breeding companies they expect to have
profits or a lot of profit immediately.
This is like a very continuous
very continuous activity.
At least breeding is based on linear
model theory. Nothing has changed there.
But AI is is having an is having a lot
of influence in
especially phenotyping.
Think of
satellite imaging and also some new
animal industries such as aquaculture,
where everything is monitored and
analyzed with very complex algorithms in
order to provide the animals with the
the most profitable environment and and
genetics.
And also we should be aware that we are
under public scrutiny the more and more
because of environmental concerns, also
welfare concerns.
But traditionally it was only because of
the impact potential impact that
breeding was having. But now it's also
the impact of AI itself. All the huge
demand on
electricity, water for the big this big
data data centers and so on. So, there
is a lot of
a lot of corners or issues that we need
to to consider here. So, so why all the
all the fuss? Why are we talking now
about AI? Because AI is is is pretty
old. I mean, already in the 1950s
even deep learning resurrection, the the
back propagation algorithm was already
published in 1986
by Geoff Hinton who was Nobel Prize a
couple of years ago. But why? Why Why is
that? So so my my interpretation
Okay, so my interpretation is that new
large language models
are actually uh
generating language and text
as if we could talk this with them as we
if we could they could understand
me and I could understand them. So
they're making very very uh sensible
uh language and text output. And we have
always been told that
if there is something that really
separates us from the rest of uh
animal kingdom is actually the language.
The language is really uh
a unique language property or so we
thought.
But the fact that then this is no longer
the case, this is disturbing. So this is
of course something that I think is
making appealing the people or or
disturbing the people or making people
that this time is is really different.
Of course, it's also different because
now we can manipulate and generate any
kind of data, protein structure, DNA
sequence, music, text, etc.
I I and I would like to argue that
mostly most of the advantage is is is
because of huge brute force, not because
now we have very very different
algorithms. I will I will show you just
in a in a second. So for instance, in
this slide you could you can see
on the left I I think you can see my my
pointer. Sorry, on the right
you can have the efficiency of
new GPUs, which are increasing
dramatically more and more teraflops
per per $1,000 for instance.
But also another interesting thing is
that AI has become a big industry. So,
if a GPU processor maybe 10 years ago
would cost
a few thousand to thousand dollars, now
they can cost up to maybe
near $50,000 or so. So, this is not
something that a normal academic group
can buy. It is something that is more
accessible only to to the big
to the big industry, sorry.
But now, what has happened also it what
has happened is that there's a big
inflation in the cost of all these um
hardware, especially for solid disk, for
for a storage. So, it's no longer than
GPU um computer cost is going lower in
cost all the time. It It is actually
going up, and this is something that
maybe you have realized already
when you were wanted to buy or update
your your your clusters on for instance.
And here on the bottom side on the
right, you can see the number of
parameters that this parameter that
these models are having, and this is
really huge.
Really, really impressive. So, here
and then well, this is for the poor
video gamers. Now, they need to pay much
more for what they used to do a few
years ago, but
anyways.
Indirect cost.
Anyways,
this is in this graph
>> [snorts]
>> you can see the number of parameters
that most that many modern large
language models have. So, for instance,
for Claude,
you they have up to 5 billion 5,000
thousand thousand thousand parameters.
So, 5 to the to the 12th parameter. So,
what is this? Well, it's a big number,
but how big?
So, just consider that Wikipedia
contains about
say five five five billion words and
about seven million articles. And
Wikipedia, we can think that it contains
basically all
all knowledge, most a big a very large
part of the knowledge that now now we
have on basically any topic, even
biographies of of people,
history,
science,
everything.
And this is about 1,000 parameters per
Wikipedia word, but about 1 million
parameters per article. So, think of the
standard methods like blob or
regression. I mean, they have one, two,
three parameters and with that we were
able to construct very
I mean, able to predict and to construct
quite reasonable models. So, now we are
talking on a scale that is 1 million 1
million larger than that. So, basically,
what these models are doing are
memorizing memorizing the whole
knowledge that is around and then
digesting it and
giving a prompt that is translated into
a numbers
writing
a reasonable prompt. But, I don't know
if you know
uh this absurd theater called the the
lesson from Ionesco
a Romanian
French
theater sorry, writer.
So, in this novel in this theater, what
happens is that the the the the the
teacher is teaching
a person who doesn't know anything. She
she doesn't understand anything. So,
what she does is to learn by heart all
possible multiplications in the world.
So,
she asked him the professor asked, how
much how much is uh
3,330
by 437.60?"
And
she responds, but she responds because
she has learned by heart everything. So,
large language models to an extent, they
are a little bit like that. I mean, they
have learned they have copied the whole
knowledge that they have been trained on
for instance about 4 months.
And then they are able to kind of
replicate, to massage the knowledge and
provide a new
a rewriting of what is already there.
So, this is something that we should
bear bear in mind.
And of course, you are more or less
aware of deep learning. Here, what I
would like to um
to insist is that DL algorithms are not
plug and play. So, what does this mean?
This means that uh you need
some is not they are not easily
reproducible because they are not like
uh
in in general accomplished packages.
They are just something that are is
staying there and on the web and
is very often
uh that the libraries are outdated. So,
you can no longer run what has been
proposed just a few years few years ago.
And because this system is
indeterminate, this has another um
com
consequence. So, the main consequence is
that um
in terms of interpretation, if we have
many different set of parameters that
can produce the same answer, so this
means that they are really
impossible to interpret because we have
different interpretations. Every set of
parameters values of parameters would
result in different interpretations. So,
this is also something we should bear
bear in mind.
And
why uh how do these algorithms work? So,
what computers of course only only
understand numbers.
I mean, we know very well
how images are converted into numbers
and and vice versa doing the intensity
the pixel intensity is a number between
0 and 255.
But how text are converted into numbers?
So, this is something that is
at least to me is relatively relatively
new. How How can the Wikipedia be
translated
into arrays of numbers and how can be
retrieved and so on.
So, the the basic idea
again is very simple. So, basically what
the algorithms do is they separate the
text into words into discrete units.
This is called tokenization and then
they are embedding embed. So, what is
embedding? Embedding is an algorithm or
is a many different algorithms that
basically they transform a vector a
sorry a text into a vector
of into an array of n dimension
whatever. And how is this done?
So, basically this is done by looking at
the words or sentences that co-occur
across massive data sets. So, across all
Wikipedia how many times linear or the
linear word occurs together with
regression word and so on. And again
across massive data sets with massive
number of parameters. So, at the end
this is how these algorithms work.
Basically by
correlations so to speak.
The interesting interesting aspect of
this encoding
is that they can be interpreted
semantically. So, for instance, it's
very famous uh
the anecdote that
if you have in a in an N-dimensional
space and you plot the coordinates of
So, for instance, Madrid, Spain, France,
and Paris,
they are close in the in this
N-dimensional space. The word Madrid is
close to the word Spain minus France
plus Paris or
other way around. Spain minus Madrid
would have similar coordinates to the
word
France minus Paris. So, what I mean is
that
these coordinates have a meaningful
semantic interpretation. So, when you
are making a question to ChatGPT or
whatever, this question is translated
into a vector, into numbers.
The The model looks for the most similar
uh responses to in terms of these
coordinates, numeric coordinates, and
transform the obtained coordinates in
back into text.
Of course, it's very relatively simple
to say, but of course, there are many
many different
>> [snorts and clears throat]
>> say nuances, there are many
difficulties, but what I want to stress
is that it is mostly brute force, brute
force followed by some intelligent
algorithms, of course.
Anyways, let's go or let's start a bit
more concisely working on um breeding.
As you can imagine, the most influence
the most important aspect of AI
initially was to in try to improve
standard linear models. So, many people,
including myself and colleagues, were
working with trying to improve genomic
prediction, blup, gblup, and so on.
And
as you can imagine, there is a lot of uh
people working on these areas. So, for
instance, recently there have been
already about 500 papers on these on
these area.
Earlier work by myself, by
Montesinos Lopez all in C mid
some other Chinese groups,
we did not find any dramatic improvement
on predictability
of
convolutional neural networks and so on.
And more recent work, I must confess
that is is quite difficult to interpret
with me
by me because the the the algorithms
have become so complicated that they're
very
even difficult to follow.
The papers and the algorithms and
they're also very difficult to to
replicate. It is not a straightforward,
but I suspect I always suspect when
there is a some new algorithm that is
really
making it much much better than a
standard linear linear models. So,
my my feeling there, I am a bit
agnostic. I am not sure that AI
algorithms would really improve upon
genomic prediction
a standard genomic prediction algorithms
like like G block or and so on. Well, a
bit of uh
uh
commercial just in case you want to take
a look at that.
And in the in the one of the first works
that we we did on the topic
on the UK Biobank, this is some of the
most important or summary results. So,
you can see here different colors
corresponds to different methods. So,
for instance, black and grey are linear
models. This is prediction accuracy,
okay? With different models and
different,
uh, scenarios. So, black black and grey
are linear models. Grey are kind of,
MLPs, multi-layer perceptron algorithms,
the oldest algorithms. And pink and
magenta are convolutional neural
networks. We did not test transformers
at the time because they were not been
invented yet, but
to me, the main message is that there
are no very big uh, differences in terms
of predictive accuracy. And I am, as I
was commenting, I am not really
convinced that this is
this is still so today. Nevertheless,
I don't think that maybe algorithms need
to really improve upon what is,
uh, standard models [clears throat] now.
They do not need to improve.
What I think would be the future would
be to combine these,
uh, algorithms
with,
uh, new data information, like climate
information,
uh,
satellite images,
uh, growth, longitudinal data, and so
on. So, maybe the idea is not to to
improve upon GBLUP, but rather to
integrate GBLUP into a wider algorithm,
like an agent or something like that. I
really don't know what how can it be in
the future, but rather to integrate all
sources or many more sources of
information. So, that that that can be a
way to include, uh, AI into, into, into
prediction.
Okay, so this is
as far as as it goes with prediction,
genomic prediction.
How about interpretability? So,
if you have been reading on AI, it has
been only highly debatable, uh,
topic. I mean, even with BLUP, deep
blood, it's always been considered a a
black box. So, imagine now these
deep learning algorithms that are even
blacker blacker and blacker.
So, interpretability itself is a concept
that is has not a single definition. It
has many different metrics and so on.
So, one possibility is to say, "Okay, so
what? I mean, I just don't care what the
algorithm is doing. I just
make a question to chat GPT. It
responds. It gives me a reasonable
response. So, I don't care whether the
algorithm is intelligent or not. I just
wants a response." So, the same thing
could go for for genomic prediction or a
strategy for a breeding company, etc. I
don't know how the algorithms
would make use of the information that
I'm I am provided providing,
but
it seems that the algorithm provides a
sensible answer and it is helping me
to
to reach to reach a conclusion
to act
or to react to that.
But, of course, we can think also, well,
if we don't know what the algorithm is
doing, if we don't know
what are the main reasons why why are we
getting this response to this question,
then we will maybe we will never be able
to go over over the
the state of the art. I mean, we will
never be able to to improve our
knowledge to understand what
we are doing or how can we what would be
the impact of this mutation on what
on what
phenotype and so on. Why why would this
happen? So, this is
really a very interesting and
and as I say, it doesn't have a single
in my opinion single response. I would
like simply to to show you
some experiments that we were doing with
transformers.
So, in a genomic prediction
uh scenario,
uh transformers compute um
the attention. And the attention scores
is more or less can be interpreted of
how much is important one is need in
relation to the prediction provided by
another snip. So, it is vaguely it is
not exactly, but is vaguely related
to the genomic relationship matrix. So,
it's a square matrix that really tells
me how redundant is this information,
how helpful is information from one is
snip
given another snip, and so on.
So,
what we did is that um
I run a small simulation to see where
the attention is related to uh P value
in a GWAS. Okay.
And you will see that
well, it is not a straightforward
interpretation.
So, here on this panel, you can see uh
the P values
of a given
uh GWAS study simulated data where we
have over 1,000 is snips, and I
simulated a single QTL right here.
So, GWAS really pick up the position,
and there is some other uh noise here.
So, this could be a false positive, and
so on.
This plot, the next plot here is what is
the attention score for all these snips
average over over layers, individuals,
and so on
for the same data set. So, it it is very
interesting, in my opinion, because
actually the attention of this is is
really much much higher than the rest of
uh is attentions
um
due to other is nibs.
So, it would seem that attention is
really actually could be interpreted as
a something similar to P value.
And you is is funny because even you
have some fake attentions here
coinciding with the fake
uh false uh P values.
And this is when we had a very very
simple model. So, we have only one
attention
a head and one layer. So, a very simple
transformer architecture. So, what
happens if I add more heads, so more
attentions within the single layer? So,
for instance, here we have
four heads, so a different model, a
different transformer model.
And suddenly you see that instead of
being a a flat
a flat line, it becomes much much uh
shaggy, so to speak. And if I add more
layers, then
things become completely un
uninterpretable. So, this is still
uh coming out, but not very not very uh
clear this. So, you can see that
attention the the very definition of
attention or interpretation of attention
would depend exactly or is very very
much on the model that we are using. And
these models can have even very similar
predictive accuracies. It's not that
they're completely different in how they
they behave.
So, this is the same thing for the QTL.
And with real data,
this is with real data, a GWAS P value
and with the transformer. Of course,
we don't we don't know in
with real data which are the true QTL.
So, here the dashes are simply
the [clears throat] peaks that are over
10 to the minus 10 to the minus three.
And we see that the profile is very
rough, but also that attention
I mean, I wouldn't say that I could
interpret this that this is
Well, yes, this is a peak here, but
not really
not really a clear clear interpretation.
So, my what I what I I guess what I
think is that
uh at best, transformer annotations can
be interpreted only in very simple
architectures. So, this is something
that I think we should bear in mind when
when we think about
these architectures.
So, the third topic that I would like to
discuss very briefly is generative AI.
So, generative AI is really
like like the queen of AI. Everybody
thinks or we think that is really an
impressive tool and and I agree.
And we all think that it's going to
solve many of our problems. I think
I am convinced actually that it has very
many
interesting and nice nice applications.
So, for instance, a generative AI is is
is a model
is is like a simulation model. It's like
a genetic simulation, but instead of
having a set of parameters and a set of
quantitative trait, etc. with given
heritability, what really these models
are doing are imitating. So, you given
millions of photographs of famous people
for instance, and the algorithm would
generate
new new photographs. And the algorithms
are improving that much that you cannot
distinguish often what is a which is an
actual image from an image generated by
computer. And you can do that with
videos, with music, images, pictures,
and so on. So, it's really having having
a
having an impact. So, there are many
different algorithms. So, like for
instance, generative adversarial
networks, diffusion models are the most
popular models
uh these days.
Autoregressive methods, flow-based, etc.
So, there is a lot of literature on on
the topic. I would say that the most
important
um
topic or or or
or
variant of generative algorithms are
conditional generative algorithms. So,
algorithms that generate do not generate
simply
any
any kind of distribution, but that
conditional on on a given feature. So,
for instance, we would be interested
in [clears throat] generating
images given DNA, for instance, images
of fruits or and so on. So, not not
simply replicating what we observe, but
conditioning on um
on a given on a given feature. So, for
instance, at CRG, there are people
working on how to generate uh new
proteins that have
improved uh features. So, for instance,
improved
uh characteristics that they bind better
to a given ligand, and so on. So, not
not simply replicating known proteins,
but designing new proteins that have a
given given uh characteristics of of of
interest.
So, these are called, as you can see
then, uh conditional generative models.
So, we applied some of these tools but
in a very or relatively old algorithm to
predict
uh fruit shapes from DNA. So, this would
be one example of application. So, we
took
This is using an old technology and I
would now we would do it much more
realistically uh with much more um
accuracy, but simply to to show you what
that what can you do with this this
models. So, for instance, what we had is
different picture from from tomato with
their crosses and some SNP uh data sets
um
um characteristics.
So, we trained the model on the
features, the characteristic and SNPs
and so on and the images
one by one by the by the other.
And then we predicted
as as you can see here they are
simply the profile. Now, we would do it
with a complete picture, but uh 3 years
ago or 4 years ago it was with uh I mean
not not so perfect, but what you can see
here is that the algorithms are able to
predict This is removing. This is not
used in the training, but when we give
it to the training
and this is what what what was
predicted.
And you can see that there is a close
similarity between the shapes of the
tomato predicted and and observed. So,
in this one here, I don't know if you
can see the the mouse, but this tomato
strain has a lot of variety variability
in the size. This is the core de boeuf,
the ox heart
variety [clears throat] and there are
there were tomatoes of very different
sizes. And what is interesting also is
that you can go beyond DNA. You don't
need to to look at DNA. So, you can
I have images of crosses. So, for
instance, this is one parent and another
parent and this is the observed
offspring. So, when we train the
algorithm, of course, without using the
the offspring or this particular
crosses, we are able to predict the the
offspring given these parent shapes. So,
this means that basically shapes of
tomato in these cases, they are they
behave kind of
additively.
All right, there are also other kind of
algorithms like variational
autoencoders. So, in this case,
what you what we you can see here is the
different strawberries
simulated with
with a random well, kind of
random shapes of strawberries. When we
are using these kind of images as as
training. So, the amount of
possibilities of the number of
possibilities is is very
very large and actually
sorry,
generative algorithms are quite
is a heavy area of of research.
>> [snorts]
>> So,
as I was saying initially, can we think
that AI will
would supersede? I mean, will replace
breeding, even statistics, quantitative
genetics?
Probably, you would say no. I mean, at
least that would be my my my intuition.
And this is what I still think. It it is
going to be very very very difficult.
The same way that books have not been
replaced by electronic means, but rather
they
cohabit. They they they coincide.
But nevertheless, in the last few years.
Uh we have seen that
quantitative genetics and breeding has
become more and more computer-based and
less and less on theory. So, if this
trend continues,
uh we can see that the role of AI
will really uh trying will be much much
higher than is than is now.
And I will show you an example to see
to see
what do you think? So, uh
we have seen the change already. I mean,
we have seen that for instance,
principal component analysis has been
replaced by UMAP when using
um
single-cell data. Why has this happened
in just a few years? So, in just the
three years, people has started to do
single-cell sequencing and instead of
using PCA to principal component
analysis to replace
uh
to visualize, sorry, the uh the data,
people has started using UMAP. And UMAP
is really
uh well, also
it's a visualization technique um
proposed by Geoff Hinton among others.
And yes, I agree that it produces very
nice very nice plot and painting and so
on, but what people don't know or or at
least most people don't know is that
UMAP depends on very many uh
on different parameters and that the
choice of parameters can have a huge
impact on what you visualize. And
whereas principal component analysis
does not have does not depend on any on
any parameter, well, only on how many
eigen values and eigen eigen vectors you
want to to plot.
But
uh
UMAP
depends
on several parameters. So, the main ones
are the number of neighbors, which
represent
how the method is going to attend or to
focus on the local versus
uh global structures. So, here you have
the same data set represented with
different values
for the UMAP
number of neighbors.
So, you can see that these are the same
data exactly as this one or this one.
So, you can
if we go back just a second, so you can
simply realize that this plot that you
you see here depending on the parameters
that you are using,
then
you will find
completely different um
completely different uh
portraits of of the data. So, what
happens that people will choose the
parameters according to what they expect
to see. So, if these are
So, for instance, if things are introns
and these are exons and these are I
don't know
uh transposon, UTR, whatever, people
would choose the parameters when they
provide
an arrangement that they like it. So,
this is really really dangerous, I
think, uh
in this area. And another parameter is
the minimum distance. So, again, this uh
the the imp-
the impact of these uh parameters is is
going to be uh
very large on what we observe. So, this
is a method based on
deep learning, basically,
that we don't understand how it works,
that has become extremely popular that
everybody is using
and that
um
none of us perhaps understand very well
what it does. And none of us think of
which would be the the parameters that
that set that would be most more
appropriate. Maybe there is none, of
course. Maybe that can that can happen,
of course.
Okay.
So, just to
to give my my final thoughts on on
on this topic. Well, as you can see,
there are many many different aspects
that could be discussed more more in
depth.
In my opinion, the the way I see it is
that
>> [clears throat]
>> at least for the industry, there are
many because they have the data and they
have the power and and they have
many people capable of
of running these algorithms. There are
many different
opportunities to opportunities. So, for
instance, one one application that I
might think is developing
different A I I agents, different
algorithms, so to speak, start
that they start to compete in each other
as they do in general generative
adversarial networks, for instance,
to guide to guide breeding policies. So,
they can they can think develop they can
really
they could develop algorithms that they
they try to mimic what would be the best
breeding strategy
and see what would be
the main one, but not deterministic.
It's rather when you have different chat
GPTs talking between between them
and so on.
Of course, another very important
application are local
large language models that they use only
internal information, so that you can
retrieve easily information and you can
store information and retrieve it easily
and in a mini mini meaningful meaningful
way.
There are There are also something
called world models, which are models
that basically try to replicate as
realistically as possible a given
a given a scenario, given a field. So,
for instance, a field in a given or a
series of fields or planting fields or
they would be like like twin like twin
models more or less.
So, there are many different
opportunities.
In academia, I I mean in terms of
research, I think that we are
trying to understand what these models
do.
They're trying to interpret
to improve perhaps some of the
algorithms and so on, but it's still I
think
we are kind of drowned by so many so
many new algorithms that appear every
day. So, it's difficult really to keep
up with uh
with what has been what is being
produced.
Well, as I mentioned, precisely one of
the most important areas, I don't think
that AI or deep learning has produced
any real improvement upon upon
traditional breeding or predictive
algorithms like GBLUP uh and so on.
But, I think that deep learning would be
very very useful if we can unify
highly heterogeneous data and be able to
produce
to produce an improved or to or to or to
to produce
a synthetic a synthetic prediction that
takes care of all all aspects of the
environment and the genetic and the
genetic
information.
Now,
a little bit on the risk of the very big
risk that I think are associated with
AI,
not
especially for for inbreeding, but in
more generally.
One important thing is reproducibility.
Science in general is based upon
the fact that reproducibility is
something given or is something that is
important for the advancement of of
knowledge. Now,
I think
this is at the stake because algorithms
are so
Someone is talking there or
Okay.
Or um
I think
this is something really really serious.
Another important aspect that we have
seen already is what would be the impact
on human learning. This is something
that I cannot
it cannot escape from from my mind
because we have
all of us have become much more lazy for
programming for instance.
Even if we find the the slightest
question, we just go to to cloud, to
whatever language model and we ask,
"Okay, give me a code to do this plot
with these colors and with this data
frame and so on."
But my question or my my my concern
is
whether this will go beyond that. I
mean, whether because
as I as I was mentioning, these models
what they do is basically they memorize
the whole world knowledge
will make us
will kind of
compromise our progress beyond what is
stated.
I don't know. As I mentioned,
mathematicians are using
these models to prove new theorems and
to find even to find new theorems. So,
this may not be the case, but
there is there is an issue.
And of course is having a huge impact on
writing and reviewing. I don't know in
reviewing how how the impact would be. I
have never used it
for writing reviews and and so on. I
cannot think that this will replace
this part of aspect. But of course as I
writing
well, we all know it has really
dramatically improved in in this case.
Science writing.
So, my general conclusion
would be it's it's a very nice time for
action. I mean, if you want responses to
a specific questions,
they are there. I mean, really
meaningful questions, responses and and
so on.
But for understanding how we go to how
we obtain that question that sorry that
that answer maybe
is is so not so
not so we we live in a pragmatical
rather than theoretical say era. So,
that would be more or less my my the way
I see the way I see the impact of
artificial intelligence modern
artificial intelligence on our on our
field, not only
as you can see in breeding, but
basically in any aspects of of science
or even the the society.
So,
well, with that I I am done. If you have
questions, you can
do it through the chat or directly. If I
Whatever you
you prefer.
>> Okay, thank you so much Miguel for this
talk that for sure gave us some
food for thought.
Uh
we have three questions in the chat and
we also are running a little bit
late on time because we have to finish
finish at 5:00 sharp for the the seminar
of the Fung Task Force.
So, the first question is by David
McHugh.
So, David, if you want, you can unmute
yourself. Or other otherwise, I will
read it from the the chat.
>> Oh, sorry. Thanks, Francesca. Can Can
you hear me okay?
>> Yeah, we we hear you.
>> Can I
Okay, yeah. Yeah,
I was just wondering what Miguel thought
about
um
you know, incorporating other omics data
types, but obviously, it would be very
useful to have those generated for the
same animals that you would usually be
using for genomic prediction and
ultimately genomic selection, but that's
not really feasible. But if you have
data from smaller
smaller numbers of animals in
appropriate experimental contexts, um
you know, gene analyses, functional
genomics experiments,
using the appropriate tissues that might
relate, for example, to meat production,
you know, muscle tissue or whatever.
Um you know, can you see, you know, how
will AI-based approaches, AI models,
really help in terms of using that data
in a meaningful way? What What does
Miguel think about that?
>> Yeah, okay. Yeah, I I I
I think that using multi-omics data
would be very similar or similar
conceptually
to what uh so for instance
um
climate data, environmental data, soil
data, and so on. The The difference
I see it as that many omics data, they
are measured in a few animals or few
individuals, whereas genomic data is
available for everybody, not to mention
phenotypic standard phenotypic
data. So for that, I think is is is
really
something that is not only related to to
AI AI. So how could AI
I mean, to me the way is could AI impute
the missing data? So that could be one
one one question. I mean, that would be
I think the
the main question. So that I I
I I I cannot
I I cannot
I cannot say. I think is is whether if
the data are there, they are easy to to
integrate. The way that I think this
deep learning algorithms have not
thought very much is about what happens
with missing data because with missing
data, normally you you throw away the
whole row. If you don't have all
information, they throw whole row. But
it of course we need we need methods I
think to to predict what is missing
there. Yeah, so Mhm.
>> O- Okay. Tha- Thanks, Miguel.
>> Yeah, thank
>> Okay, so the next question is is from
Edgar Caballe- Caballero Vargas.
So Edgar, I already see that you unmuted
yourself. So
>> Yeah, so thank you.
Yeah, thank you. Thank you very much for
this talk, Dr. Perez. So my question is
basically like with conditional
generative AI, I saw that you basically
could generate the tomato likely
phenotype. So I was wondering like
related to the previous question like if
you can integrate different like
multi-omics data it could be like if you
want to like know what it's this like
like I want to generate a genotype, for
example, that could have a potential.
What type of like conditions could be
could be like phenotype data or another
type of I don't know multi-omics? Is
that feasible?
>> Yeah, yeah, that that's a very very very
interesting project and actually sorry,
question.
And actually we we have a project we
would like to explore
these issues. There is some software,
for instance, Evo3, I think, that is
already able to generate complete DNA
sequences but from prokaryotes
individuals. But one possibility would
be to use generative algorithms to
generate the whole for instance, a SNP
data and the whole associated um
features. So, for instance, the the
shapes of fruits, root architecture, or
the conformation, color patterns from in
dairy cows, for instance.
Um yeah, that that I think would be
feasible and and kind kind of doable.
Um the question is to have the data also
to be able to train the to train the
algorithm. But it's something that in
principle should be should be should be
doable. Maybe not now, but maybe in in a
few I don't know in
in some time in in the future.
Um yes,
one thing that we were discussing in a
paper with with Gustavo and Laura
Fingaretti was how to combine also
standard simulation with generative
with the generative algorithms. And I
think that they they can they can
complement.
Oh, oh, and you mean you mean also to
generate or to combine all data. Well,
in principle, you could you could
simulate also omic data. That So, for
instance, is very relatively simple to
generate
abundant data, microbial abundant data,
in a few lines of code omic data,
inspection data, and so on. Maybe is a
bit more tricky, but uh
Yeah.
>> Okay, so the last question for today is
from Andres Segarra. Andres, do you want
to
unmute yourself or should I read it?
>> Hello, Miguel.
>> Hey. Hi, Andres.
>> I've been told that there's been a
proliferation of
grants with the use of AI.
And the papers that I read that deal
with AI, they are often full of cryptic
models they don't understand and
vague statements.
And my impression is that essentially we
are kind of back to 1960,
where instead of reading papers, you
just ask your friend
or trust your experience.
And this has to do with your last slides
that you look you are very worried about
the state the advancement of scientific
knowledge.
Yeah.
>> Yeah, I I heard that So, for instance,
they were an inflation of
ERC grants and they were not able to
review all of them. So, that's that's
Yeah, that's really something that um
concerns concerns me and and concerns
us, I think. Uh
whether
what is the the
what you see is is is is make sense or
or is is is
is credible or not. So, that's that's
that's something that I still we don't
know we don't have the the answer for
that. But uh Uh,
there was something who was saying,
"No." These days you is not that I see
It is not that I believe what I see, is
that I see what I believe. So, when you
have a pre-conditional
thought or say an idea or could be even
a grand proposal
all of all of your effort is is going
through through that. You only look at
the data or at the
uh, what supports your your your
your your your opinion, so to speak.
I don't know if you were
talking about that, Andres, or um
>> Yeah, more or less.
>> More or less.
>> [laughter]
>> So, whether I trust my friend or my own
experience. So, well, your own
experience depends on your age.
Um
Yes, um I think
that
uh you will never have enough enough
experience unless you you work in a very
reduced area of
of research. So, I think a balance you
need a balance between
uh your friends and and experience and
and new knowledge from language models
and so on.
>> Okay, thank you so much again, Miguel.
Uh, I hand over the the talk to Jose
that we close it and thank you all for
being here.
>> Yes.
>> All right.
Thank you very much for the nice
presentation and the beautiful
discussion. Uh, I just want to announce
that next seminar will be announced by
the usual channels. Uh, and as always
will be the third week of of July. We
are working to confirm the speaker.
And before we close, I also wanted to
quickly mention that right after our
session at 5:00 p.m. European time,
there is an another seminar
uh, that may be very interesting also
for you. This is from the farm variant
function task force that is kicking
kicking off today and it's about
methods.
Today is the first session and and the
speakers are Fin Gray from the
University of Edinburgh who will discuss
how to bring high throughput phenotypic
screens to livestock species and Shiva
Abdsec from Google DeepMind who style
will cover how to advance regulatory
variant effect.