Alberto Perez: AI and Structural Biology From Protein Structure to Molecular Design (TSVP Talk)
Watch on YouTubeVideo summary
Alberto Perez opens his talk by introducing himself as a computational structural biologist whose work focuses on understanding how biomolecules function and interact within the complex environment of a cell. He explains that while experimental techniques like X-ray crystallography, NMR spectroscopy, and cryo-electron microscopy have revolutionized our ability to visualize protein structures, these methods are often expensive and limited in scope. Consequently, there is a growing need for computational approaches to predict how proteins fold from their amino acid sequences into specific three-dimensional shapes. Perez illustrates this process using the analogy of building with Legos, where 20 different types of amino acids interact based on chemical properties—such as charge and hydrophobicity—to form stable structures like alpha helices and beta sheets that perform essential biological functions.
The core narrative shifts to how artificial intelligence has recently solved long-standing challenges in this field, particularly the protein structure prediction problem which had persisted for decades. Perez recounts the history of blind prediction competitions where early attempts failed miserably until a massive influx of genomic data transformed the landscape from physics-based modeling to AI-driven solutions. He draws parallels between IBM's Deep Blue defeating chess champions and Google DeepMind's AlphaGo beating Go masters, explaining how these breakthroughs paved the way for AlphaFold. This technology achieved unprecedented accuracy in predicting protein structures, earning half of the 2024 Nobel Prize. However, Perez emphasizes that knowing a static structure is insufficient; true understanding requires grasping molecular recognition—the dynamic interactions between proteins and other molecules—which involves concepts like lock-and-key binding or induced fit where flexible peptides change shape upon interaction.
To address limitations in AI models regarding predicting which specific compounds bind best among thousands of options, Perez describes his lab's innovative hybrid approach that combines the speed of deep learning with the interpretability of physics-based simulations. While standard diffusion models can generate vast numbers of protein designs or predict binding events quickly, they often lack confidence in ranking them by affinity without explicit physical constraints. By integrating AI predictions with rigorous physics-based energy calculations, researchers can validate results and filter out false positives, significantly accelerating drug discovery and enzyme design processes. This synergy allows scientists to screen entire libraries of compounds efficiently while maintaining the scientific rigor needed for experimental validation, effectively democratizing protein design so that labs worldwide can now create novel enzymes or therapeutic agents in days rather than years.
The presentation concludes with a discussion on inverse problems where researchers use AI to design proteins from scratch based on desired functions, such as creating mini-proteins that bind to cancer antigens to enhance T-cell therapy. Perez highlights the importance of distinguishing between mere prediction and true mechanistic understanding, noting that while we can now predict structures accurately, explaining *how* a protein folds so rapidly or how cryptic pockets emerge remains an open challenge involving complex energy landscapes and kinetic pathways. He argues that future advancements must move beyond static snapshots to understand dynamic processes like conformational changes required for function, such as the flapping wings of ATP synthase. Ultimately, Perez envisions a future where combining AI efficiency with physical principles will unlock new capabilities in biotechnology, allowing us to design molecules that nature never evolved and solve complex medical challenges through precise molecular engineering.
Read the full video transcript
Hello. Hello everybody. Thank you very
much for coming
um to today TSVP talk. So today we have
Alberto Perez
uh whom I'll introduce briefly. Um so
Alberto did his undergrad and PhD in
Barcelona.
Then he moved to the US to work on the
protein folding problem with Ken Dil.
And then since 2008
um he leads his own research group at
the University of Florida currently a
associate professor and associate
director of the quantum theory project.
Um and today he will be talking about
the AI in the context of structural
biology. So yeah thank you all for
coming.
>> Hi everyone and thank you for being
here. Uh when I asked who will come to
the talk I wasn't sure if anyone would
show up. So, it's really energizing to
see so many people here. Um, I'm here
for three months. I've stayed already
six weeks. So, I have six more six more
weeks to go. So, if you like any of
these, please come talk to me
afterwards. Um, what I thought I would
do today is give you an overview of my
field, the kind of things we do. So, I'm
a computational stctile biologist. I
work in trying to understand
biomolelecules, how they work, how they
interact with each other. And in our
field, as in many other fields, AI has
made uh a huge revolution. So I'm trying
I'm going to try to tell you how we're
going from protein structure prediction
to molecular design. And one of the
things that I came here to Okinawa to do
is to try to understand better how
enzymes work, how they are able to
perform their work. How can we design
novel enzymes that work better and do
things that we desire them to do? And so
the big idea about proteins and
biomolelecules is that um there are many
of them of many different types. This is
an artistic representation of what might
be going on inside one of our cells. So
you have multiple different kind of
proteins. They are diffusing around.
They're interacting amongst themselves
with other molecules membranes. There
are some small molecules, there are some
large molecules, molecular motors and
whatnot. And so the big idea is that
inside of our cell there's a lot of
activity going on. We have a lot of uh
these biomolelecules, proteins and
whatnot that in order to satisfy their
functional role they need to interact
with each other. And so these molecules
they are small. So if this is a cell
proteins are about 5 to 10 nanometers.
We cannot see them with our naked eye.
They interact with all kind of
molecules. Some of them like aspirin are
much much smaller. And in order to for
us to understand how these molecules
work, we need to know what they look
like. So knowing what they look like
means understanding where every atom is
placed. And to do that, we have some
experimental techniques such as X-ray
crystalallography. So we try to
crystallize proteins, bombard them with
X-rays, and then where they defract that
must be the places that they collide
with the nucle. So we can backmap what
the stack of a protein looks like. We
can also use NMR. So trying to use
nuclear magnetic resonance in order to
understand what atoms might be close to
each other and traced out what the
effect of the protein is. And finally we
can use carry. So since the resolution
revolution there has been like a massive
amount of factors. There's a lot of
expertise here at OIS about carry and so
those are the three main techniques by
which we can know how proteins look
like. The challenge is that these three
techniques they are expensive they might
not be amenable to every protein and so
there there might be other ways in which
we can predict or try to understand what
the structures of these proteins look
like. And so for someone that as a kid
like to play with Legos, proteins are
the ultimate Lego. So they are made of
20 different building blocks amino acids
and the way that they're made is like a
bits on a string. So you're putting
different amino acids one after another.
Each amino acids. So each different
color represents a different kind of
amino acid. You can divide them in four
major classes. So blue are positively
charged amino acids, reds are negatively
charged amino acids, green are polar,
and white are hydrophobic. They like to
shy away from water. And so under the
influence of the environment and the
properties of these amino acids, some of
these protein sequences can undergo a
transformation. They can evolve. And so
I'm going to represent this string of
amino acids with this string. And I'm
going to highlight here four residues.
These are four hydrophobic residues.
They like to share away from the water.
I'm going to put in this string four
different uh places in yellow. I'm going
to explain that later and one of the
pieces in blue. And so what I'm going to
show you is a process that we call
process of protein folding. So by
looking at the interactions between the
different amino acids and looking at the
interactions with the environment, you
have every now and then some stoastic
kind of interactions that ultimately
lead to the folded uh protein. So at
some point those four hydrophobic
residues that sh away from the water,
they interact with each other. They
clamp together at the core of the
protein and they're protected by other
residues that are polar and interacting
on the outside. So this is a folded
protein. You see that the four different
yellow pieces they are interacting
together. They form what we call a beta
sheet. And this blue region here is
making an alpha helix. So this is how
this protein falls. You have many
different kind of proteins. They come in
all shapes and sizes and they do
different things. So you have some
proteins that are enzymes. They're able
to break and make bonds. You have
proteins that are able to recognize
other molecules maybe in disease. You
have proteins that for example allow
your muscles to contract and extend. And
you have membrane proteins that allow
things to go in and out of the cell
either molecules or signals coming out
of the cell. So one of the things that
we want is to understand what the
structures are like so that we can
understand what are the functions of
these molecules.
And so I've shown you the process of
protein folding. There are two kind of
competing interests. On one side you
have the chain entropy. So every time I
put like a something inside my pocket it
comes out like a tangled mess. So that's
the entropy. And on the other side you
have like the interactions that these uh
amino acids can make with each other
that favor folding. And so you have an
entropic term an enthalpic term and
there's a balance between the two. For
some protein sequences what you find is
that at low temperature you can find
this native state this folded state. And
this is a panel like representation
where at high temperature you find
unfolded states at low temperature you
find the folded state. Now how many
protein confirmations how many
confirmations can a single sequence
adopt? Many many many many many
in fact more number of possible uh
configurations than the number of
hydrogens in the known universe. That's
how many possible confirmations every
single sequence can do. So if a protein
had to go through every single
possibility before finding a native
state, it would take longer than the
edge of the universe for a protein to
fold. Obviously, we cannot wait that
long. And so all proteins fold very
fast. The fastest proteins can fold in
milliseconds and even in microsconds.
And so that's how how fast they fall.
And so for the longest time, this led to
some questions in in our field. It's
what's called the protein folding
problem. It's a 60 year old kind of
problem and it has three parts.
The f the first one is how do proteins
fold so fast? What are the structures
that they adopt? And third, can we use
computers to predict those structures?
And so these last two questions have
been mostly solved recently and I will
tell you how.
These third questions, how do proteins
fall so fast? That's still something
that we are discovering. So I've told
you this is a 60 year old kind of
question. So for the longest time if you
look at papers historical papers you
will see oh we solved the protein
folding problem and then someone else
publish another paper. Oh we solved the
protein folding problem. It turns out
like yeah maybe they did something with
one or two proteins and then you tried
to apply the methodology to a new
protein and it didn't work. And so one
of the biggest things, one of the
biggest contributions to our field is
the emergence of blind uh prediction uh
competitions. So these are events in
which someone is going to say well we're
going to start a competition. Everyone
in the world can participate. What we're
going to do is we're going to contact
some experimentalists that are working
on crystallizing or solving with Kimm
some of the proteins. We're going to ask
them for the sequence and then we're
going to provide this sequence to
everyone who wants to participate.
So this usually happens every other
summer. So during summer we have two
months. Every day we're given about two
sequences. The competition is going on
right now actually. And for every
sequence you have about three weeks to
predict up to five stractors.
So you submit your pipe structures and
at the end of the competition someone
who has not participated will be
assessing all of the predictions. This
is what we call a double blind
experiment blind to the predictors
because we don't know what sequences
we're going to be given. So methods uh
cannot be um geared up for a particular
sequence and then it's blind to the
assessors because they're given all the
predictions but they don't know who
predicted what. And so at the end what
you get is what is the state-of-the-art
of the field? What are the best methods
to work on the protein solving uh
protein structure prediction problem?
How do you think they did on the first
chance?
Good, bad,
bad, right? Yeah. So the first
experience was aartic experience. So if
I were to go back in time to 1995 and I
were to say well this protein has all
the atoms on top of each other I would
have one CASP that doesn't have any
physical kind of sense but according to
the metrics that they use that would
have been the best protein that's how
bad we were at predicting proteins. All
right did we get better over the the
years? Well, this is a plot that is
going to try to explain that the answer
is no. We didn't get much better. So
each line here is a different
competition. So every two years we have
a new line. On the Y ais you have the
ability to have accurate predictions. So
good means the atoms are in the right
place. Bad means
it's random noise. And so on the x-axis
I have whether the protein sequence that
we were given is easy or hard. How do
you tell if a protein is or a sequence
is easy or hard? Well, I'll tell you an
anecdote. So I came to Okinawa. I'm
originally from Spain and I was walking
around and I heard something that looks
like sounded like a duck. I thought I
went around. I looked for it. It looked
like a duck but not like any duck that I
knew about. And so I said, well, it
sounds like a duck. It looks like a
duck. It must be an Okinawan duck,
right? So they do the same for protein
sequences.
If a protein sequence looks like another
sequence for which I know the structure,
I'm going to say this sequence most
likely falls to that structure. So
that's how we define an easy sequence.
And for easy sequences, we get good
structures. For proteins that don't look
like other proteins, they're sequenced.
we have very bad ability to predict
their structures. So all right so over
20 years almost no progress. What did
change is how much data we accumulated.
So back in the 1990s we knew about
50,000 sequences. Then came all the
genomic and metagenomic projects. We now
know about over 200 million uh protein
sequences. We used to know about 4,000
structures deposited in the protein data
bank. we now know about 300,000s. So
this changed from a problem of like oh
can we use physics or can we use a
statistical mechanics to predict the
structures of proteins to an AI kind of
challenge
and so in my field computer programs and
um science have gone hand in hand. So in
the 1990s IBM developed deep blue. It
beat the reigning champion Kasparov at
the game of chess. So basically the idea
is that by brute force a computer can
calculate more moves than a human can.
At any point in time you can look at the
chess board and you can decide who's
winning. And so with that uh deep blue
beat Kasparov.
So IBM said oh now we're going to solve
protein folding. And they didn't.
So a little bit later, a company named
Deep Mind is a subsidiary subsidiary
from Google. They decided to play the
game of gold. Now, very very typical
game in in Japan. And the idea is that
here you're putting stones either white
or black and you're trying to uh capture
the pieces of the of the opposite uh
side. The problem is that there are many
more possibilities than in the game of
chess. So now you cannot guess by brute
force predict who's going to win. And
there's no way to know who's winning by
looking at the board. You really have to
play out the game. And so if we were
just waiting for a computer to be fast
enough, we would still not be able to
beat the reigning champions. And yet
AlphaGo was able to defeat the world
champions.
And so and one of the particular
interesting things so this was five
matches I think it won 4-1 and in one of
the matches the com commentators there
are commentators for these games said oh
the machine has gone crazy it has made a
move that makes no sense
100 or something turns later it turned
out that that move was critical to win
the game and so it had hallucinated a
move that no human would ever make and
Yet it was critical to win the game and
this is something that will come back
later because when AlphaGo won they said
well our next challenge are proteins
and so I've shown you that there was no
progress in 20 years alpho came into the
so alpha go changed to alpha fold and in
two years we went from here to all the
way here. At this point they were just
newcomers to the game. they were just
doing everything that everyone else was
doing. So they were just
they had more resources, they had more
people working on it. So in a few months
they were able to beat the whole field.
Two years later they said well we can
now develop new transformers that pay
attention to the sequence and structure
similarities and now
this is the state-of-the-art.
And this this technology was actually
the one that recently won the Nobel
Prize in 2024 for STO prediction. And so
um of course I was at this conference
where they announced the winners. I came
back and my lab was like oh why did I
ever join your lab like this is going to
be a disaster. I'm never going to get a
job afterwards. This is this is
horrible. And of course me being the
optimist that I am I said don't worry
guys
is just the beginning. So we have Alpha
Pole. It's a great new tool. Let's use
it to be able to do new science that we
couldn't do before. And so the big thing
is like if I if I go back to the duck
that we saw earlier, you see the duck
with the with the wings tucked in and
you never know that a D can expand its
wing and then flap them and start
flying, right? So knowing something
about the structure or one state doesn't
tell you anything about the function. If
I show you this protein called ATP
synthes,
so what? Yeah, it looks colorful. It
looks nice. What does it do? Well, this
protein is actually a motor that it
captures uh it's a hydrogen pump. So, it
captures hydrogens and then it's able to
make ATP, the molecule that uh allows
you to carry energy many of the react to
do many of the reactions in the body in
the cell. And so, it is critical to know
the structure. Yes. But that's the start
to be able to understand how these
molecules work. And one of the key
aspects of how molecules these molecules
work is the concept of molecular
recognition. So if you think about your
daily life, so this morning I was on my
bike biking here then you know right now
I'm holding this I might be holding a
pen interacting with a keyboard. So in
order to do different things we're
interacting with different objects. In
the same way, proteins in order to carry
out their work, they need to interact
with other molecules. Some proteins will
interact with nucleic acids to maybe
express information in our genes. Some
will interact with other proteins or
peptides or even small molecules like
aspirin. And so understanding molecular
recognition is key to many fields. In
biio medicine to design novel molecules
that can stop disease. In biotechnology
to design uh custommade sensors or um
or possible enzymes that can catalyze
reactors that we care about.
Biomeaterials to do for example
selfhealing. uh you put two peptides
together, they form a hydrogel that can
uh capture cells to close a wound or uh
to the sign novel functionality that we
have never seen because evolution never
thought of it. And so if I'm thinking
about molecular recognition, it goes
between two extremes. On one end, you
have something that is called lock and
key kind of interactions. So this is
where you have a molecule. it might have
like a cavity or something where a small
molecule can bind. And on the other
hand, you have what's called dynamic
recognition. So proteins that might be
intrinsically disordered, they can adopt
many shapes and once they find the right
binding partner, they adopt very well-
definfined shapes. So if I had to put a
cartoon representation, this is kind of
what I thought about. And on the other
side you have your hands which can adopt
many different confirmations depending
on what they're interacting with. Right?
So one [snorts] of the things is knowing
a structure is not enough. Now we need
to understand how this works. So for a
lot of the things that I do in my lab
like we work with physics based models.
So in physics like [snorts] you try to
put a physics model you use the
equations of motion Newton's equations
of motion you propagate them through
time and you see the wiggling and
jiggling of atoms and if you look at
this for long enough you might be able
to understand the behavior of
biomolelecules.
The only problem is that these
simulations are computationally
expensive very very expensive. So right
now about 20 to 30% of supercomputer
time is devoted to these kind of
simulations and it produces terabytes of
data and yet understanding what these
molecules do or reaching time scales
that are relevant is hard and so how do
we predict binding? So I'm going to tell
you tell you about the two different
regimes the lock and key regime. So how
do we predict binding force a small
molecule like aspirin and how do we
predict if we have many different
compounds which of them might be the one
that binds best
and so
with physics it's very hard to predict
things in absolute terms which one is
the best but if I give you something
like this horse so if I give you
something like this horse it looks fast
this horse also looks fast but if I were
to say which one of them is faster, the
best way is to do a horse race, right?
And so we do exactly the same with our
molecules. So when we have this lo and
key interaction, we compete different
compounds in order to understand which
one of them might have the best binding
affinity, bind their strongest. And so
this is something that we use routinely
in pharma. So in drug discovery, this is
the early stages of drug discovery. It's
widely used. You can process thousands
of compounds, no problem. On the other
hand, when you have flexible molecules,
now that becomes a problem because not
only you have to care about what
molecules are interacting with which
ones, you also have to care about what
are the different shapes that they might
adopt and what is the right one. And so
it's a much more complex problem.
So we've worked in this area. We've
developed some methods based on physics.
And basically the idea is that we have a
protein target. we have two possible uh
compounds. They're disordered when
they're on their own and they might
adopt like a helix like extractor when
they bind. So we can make these
comparisons. It takes us about 30 GPUs
running for five days non-stop in order
to make one of these predictions.
Now imagine that I have a 100,000 of
these compounds to test. There's no way
I can do it with a computer. So as soon
as Alpha came out, what did we do? We
said well you know they design alpha
fold to predict the structors of
proteins
but a peptide binding to a protein
that's a similar kind of problem it
might work and so people started using
alpha for all kind of things that it
wasn't designed to do. It turns out it
turns out that it did pretty well. It
did really well at predicting protein
protein protein peptide and other kind
of interactions.
It was also really good not only at
predicting structures but people started
doing things like uh flow matching and
whatnot to predict ensembles predict
different confirmations away from the
static structure. They also use it to
design novel proteins and so from alpha
fold that was designed to do only one
thing a lot of different functionality
emerged and that has opened the field to
many many different kind of things.
So still we were like okay
but we don't care only about knowing
what's where something binds. We care
about knowing which of these compounds
might be the best. And so we said if I
give you all of these little compounds
that have go from strong binders to weak
binders
will alpha know how they bind and which
ones are the strongest? So it turns out
that alphapold did very well at
predicting that all of these fine but it
could not tell us anything about which
of them was the best.
>> Can I ask a question? So these molecules
live in some environment aquas or
something else. Where is that
information in here?
>> So in the AI model that's implicit
because it's been trained on um either
NMR structures which already captured
some environment crystal graphic
environment environment. So they capture
whatever environment they were in.
Right? In a physics model, we would
explicitly put the environment, either a
membrane environment, either an aquis
environment with ions, a crowded
environment, so on and so forth. Thank
you.
>> So AI is very confident about these
different peptides should bind, but we
have no idea which one of them is best.
And so, um, if I were to ask you, do you
like Japanese snacks? If I ask you about
this snack, do you like it? Pretty
confident. Yeah. Do you like this one?
Yeah. How about all of these? I'm pretty
confident you like all of these. But
being a 100% confident on each one of
them doesn't tell me anything about
which one you would choose, right? So,
confidence is not the same as
preference. So we said well if alpha
fault is really confident about all
these predictions could we ask it
something about which one to choose and
so we started thinking so my student
leeway he started thinking about like
well maybe in this uh function that
alpha fold is encoding it has learned
something that looks like a scoring
function and could we instead of
predicting one peptide at a time could
we feed it two peptides at a time And
now it has to make a decision which one
are you more confident with this one
being in the active site or the other
one. And by doing that we could predict
which peptides fine and which ones are
the strongest. And the big game for us
is that when we use physics we take 30
GPUs running for five days. When we do
this AI approach it takes us one GPU
running for five minutes. And so now we
can scale to hundreds of thousands of
compounds. So what that means is we can
go get whole protein libraries, screen
those protein libraries,
uh cut them all into all their little
pieces and identify which ones bind,
which ones do not bind and which ones
are are the best. So we're working with
some immune response proteins. We're now
understanding many proteins that are
able to bind these immune response
proteins from viral proteins from
bacterial proteins and even host
regulatory proteins. And so this is
allowing us to to learn new many new
things. Now one of the things that we
are seeing is like well AI still has its
own information. It's fast. It's really
good but it might have some biases. So
what our new paradigm is is that we're
combining the speed of AI with the
interpretability that physics based
models are giving us and with that we're
now able to make more predictive uh
molecular modeling. So basically we have
a front end in which we use AI at the
sequence level then at the structural
level and then we go to increasing
degrees of complexity of physics-based
models. When we are done with all our
filters, we go down and we send our
predictions to our experimental
collaborators and with that we have been
able to identify novel uh proteins and
we have also been able to distinguish
some proteins that were previously
characterized and were not biologically
relevant.
All right. So so far I've told you about
protein structure prediction the problem
of molecular recognition and then the
next stage that I want to tell you the
last part of this talk is about going
from a structure prediction to design
right so the idea of design is the
inverse
problem as protein folding so in protein
folding for protein structure prediction
we have a sequence and we want to know
what is the 3D structure that it adopts
And that has been mostly solved by AI
in um in protein design. What we want to
know is if I have a 3D structure of a
protein, what is the sequence that is
best suited to adopt that structure?
And so the structure could be a
structure that it's already known or it
could be a noble structure, a structure
that is hallucinated for example by
alpha 4. And so the main idea here is
that during evolution things are
things evolve until a certain point. So
you have some parameters. For example,
most of our proteins have to be as at
least stable at 37 Celsius. when you
have a high fever, you are putting some
cold
eyes on your forehead or whatn not to
lower the temperature so your proteins
don't unfold or what not. And so with
designed proteins, you can now go to
proteins that are stable maybe 80, 90
Celsius, what not. And so one of the
biggest players in the field is David
Baker. For the longest time, designing
proteins was a task that very few groups
in the world could do. David Baker was
the maximum exponent of what we could do
and before AI they were already doing
very impressive things. So this is uh
one of the designs that they did. It's
basically a protein that is able to fold
into a 3D structor. It's able to
assemble itself with other proteins of
its type and make this cage like kind of
compound. So you have 40,000 atoms here
that when you solve the experimental
structure, those 40,000 atoms are on top
of the design. That's amazing.
So the challenge is that
only one or two groups were really good
at doing this. The success rates were
not high and it it required a lot of
manual intervention. So it was very
artisal in some way. And so the idea is
could AI help in this front as well. And
so from 2015 where Baker lab was the
maximum exponent to today what we're
seeing is that we have now democratized
um design. So basically any lab can now
within one or two days say hey I want to
design a protein that binds this
particular target. And within two days
you have 10,000 designs. Now are all
those designs good? No. You still have
to find out what are the best. The
success rates are 10 to 20%. And
programs are getting better and better.
But like having like this AI machinery
has helped a lot. So this is the Nobel
Prize. 50% went to David Baker for
design and 50% to Disabon Jumper for
sector prediction. So how does that how
does the design work? So it works pretty
much the same way as in the image
generation field. So we use what's
called diffusion models. So the idea is
that if I have a picture of a dog, it's
made out of pixels, right? I can
randomly change a few of the pixels in
this picture and I will get something
that still looks like a dog, but not
quite. I can learn the process of going
from this picture to this picture, but
also from this picture back to this
picture, right? And I can do that again
and I can do that again and again and
again until I get to a random set of
pixels. Now if I have this set of
pixels, how do I change them to get a
new image? I have no idea.
But if I have a random set of pixels,
I know how to generate another random
set of pixels.
And if I've learned the process of going
from a random set of pixels back to an
image, I can say, well, this is my new
set of random pixels.
And if I go back, oh, now I get a cat.
So, of course, to do that, you have 10
of many, many, many images. Yeah. So, I
can do exactly the same with proteins. I
have about 300,000 proteins. I know the
position of all the atoms in the
protein. I can add a little bit noise to
the positions of those atoms. And I can
learn that step. I can also learn the
step of going back. And I can go all the
way until I get a random distribution of
atoms in space. And from that random
distribution, I can generate a new
random distribution. Go back and
generate a new protein. So that's how
design works. Nowadays once I have the
fall, I calculate what is the sequence
that is best suited for this uh protein
design. And so again like we've been
using
the latest AI technology combined with
our physics based pipeline. So basically
we start we use diffusion models in
order to predict as um new proteins that
will bind a particular target of
interest. Then we do all of our AI
pipeline to filter out which we think
are the best designs. And then finally
we use our physics based models to say
like do these two things that have been
completely orthogonally uh designed do
they agree or do they differ? When they
agree we send those for experimental
validation and this has been successful
successful in many cases. I started the
talk by giving you an example about
blindness structure prediction.
There have been new challenges for
designing proteins. So this was one that
we participated a little while ago. It's
called the biome challenge. It's to
binders. So the idea is they gave us
this target. So it's this is a cancer
cell. It has a tumor antigen and usually
a T- cell will have to recognize that
antigen, bind to it, destroy the the the
cancer cell. Right? So they said, could
you design a mini protein that would
bind this antigen? And if you can, then
what we're going to do is we're going to
take that mini protein, we're going to
put it in a T- cell. So we're going to
have a functionalized T- cell. This is
called a CART cell. And so once once we
have that, what we're going to say ask
is, can those designs
be successful at killing these cancer
cells?
So basically that's what they tked us to
do. About uh 12,000 sequences from
around the world were submitted. Each
team could submit 500 sequences. So once
they gather all the sequences then they
did all these kind of functional assays.
I have no idea about these functional
assays exactly what they are. The basic
idea is they tested whether the carti
cells were able to prol proliferate in
the presence of cancer cells. If these
carti cells were able to um send
signaling molecules that would activate
um the the killing of the of the cancer
cells. And finally, what was the ability
of these T- cells uh to kill cancer
cells without killing regular cells? And
it turns out that our submission was the
top submission in the whole competition.
So it was really nice. They 3D printed
uh awards for the top teams and they
sent it to us. So this was the mini
protein that we designed and this is how
it binds its uh uh cancer antigen. And
so like this was a nice blind
competition. One of the biggest surprise
is that the best binder was not the one
that won. The best binders actually
create exhaustion and so it doesn't
allow cancer cells to be killed.
Um and so this was one of the lessons
that we learned along the way. So
by now I have told you three things. I
told you about protein structure
prediction about how that's not enough.
We need to understand molecular
recognition how these proteins interact
with uh different molecules and the hope
is that we can use that to understand
how proteins function. So basically this
challenge of protein structure
prediction is mostly done.
A lot of molecular recognition is now
what what's going on in our field like
there are many things that we cannot
still understand using AI models for
example things of like how do cryptic
pockets emerge or how does a lost happen
like those are too complex for most of
the AI models that we have right now and
the future where everyone is trying to
go is how do we uh look into protein
function how do we go from a sequence to
function and so um with that I just
wanted to finish with picture of my
group. Really happy to be here. I'll
take any questions. Thank you.
>> Sure.
>> Yeah, please if you have any questions
um just please wait for the microphone
for the people on Zoom.
>> Yeah. Thank you very much. It was so
nice to see such a good presentation. Uh
so I have two questions. The first
question is about the approach when you
combine AI and physics based approach.
So what is the lack of just using AI and
the sequence? Um what's the lag there?
What we take from the physics? Is it
just the more parameters?
Because like
physics is kind of embedded in the
sequence I would say.
as well.
>> All right. So the main thing for us, so
whenever you use an AI model, one of the
bigger biggest question that reviewers
can ask you is, oh, what about
memorization? Like did you get the right
answer because it's making a prediction
or is it because it has already seen
this and so it's giving you the right
answer for the wrong reason. If it's
giving you the wrong answer for the
wrong reason, then it's gonna make
predictions that are false positives.
when you use our physics based model
because it has it hasn't been trained on
data. It just uses like you know how do
bonds stretch, how do angles change and
so on and so forth. If the answer with
our physics model agrees with the AI
model then that increases our
confidence. We could still be wrong but
we're much more confident because they
have completely separate origins.
>> Okay. And I have the second question
also short one about the diffusion and
the protein design. So are there any
like common cases when it doesn't work
>> or just by chance?
>> Yeah. So so all right. So if I were to
say okay I'm going to use RF diffusion
to predict binders to this molecule and
then I'm going to design the best
sequences that are able to adopt this
sector. The success rate is expected to
be 10 to 20%. So that means one of one
every five will fail. Of course, you can
only test experimentally. Like my
collaborators will try at most five,
right? So if I choose five that are
incorrect, they're never going to
collaborate with me again. So we need to
make sure as best as we can that what we
are giving them is enriched from having
one every five correct to three or four
every five correct. And so that's why we
do this complicated machinery. There are
so deep mind has done a new AI that is
proprietary. So this is not available to
the I think it's called alpha proteio or
something like that. So this one is not
available but this one is much much much
better at making. So it goes from 10 to
20% to a much higher percentage of
success.
>> Yeah. My question was more about like we
know that let's say basic alpha fault
doesn't work well for uh instly
disordered proteins let's say maybe
isn't this like the case for diffusion
models as well
>> so
for disorder proteins for predicting and
designing disorder proteins I don't
think it's going to look very good
>> like I think like there are some AI
tools that are starting to predict how
to bind disorder regions.
>> Yeah.
>> But not necessarily
to make disorder region. I mean like
there are other tools that can give you
disordered regions. But usually I think
the state-of-the-art right now is I have
something that is disorder. I will use
NMR and I will do a population of states
rewe them so that they satisfy the NMR
data. That's I think the best that you
can.
>> Thank you.
Thank you Albert. That was really nice
to see and um so so this is probably a
very simple question but still you had
in the beginning the like a landscape of
energies and you you talked about the
the native state when the lowest energy
right
and then you also had one question that
was still you working on the how how
they can solve this hard problem so fast
in reality I guess I mean do do they
always find the lowest one or or is
No or I mean
>> so first not all protein sequences will
have a folded state
>> some of these proteins will have like a
more complicated landscape where like
you have like maybe a misfolded state
that can aggregate and so it they don't
find they find this aggregated state
before they find the native state. So
you have like um um what are they called
chaperons that will help fold some
proteins that are misbehaving. So some
of them need help. M
>> but eventually they will find the
absolute smallest one or
>> well and when they don't we get all kind
of diseases right like Alzheimer's
Parkinsons and what not
>> thanks
I'll be here six more weeks so I'll be
two floors above if you have questions
>> can I add one more and then
>> yeah just
>> yeah okay so the I remember when that
they got the Nobel Prize but you know
the because it's a statistical you know
method in essentially
that
>> I would not call it a statistic
>> okay but I remember the statement was
that now we understand how proteins are
folding I it's a good predictor but it's
not the explanation right
>> it's not the explanation it's a black
box still
>> so all right let me backtrack a little
bit so when I said the protein folding
problem is these three questions.
>> So for me the protein folding problem is
not solved. The protein structure
prediction problem is mostly solved.
>> Right?
>> And so depending on who you ask in the
field, they will not distinguish those
two things. For me protein folding is
understanding what are the myriad
pathways that will take you from
unfolded state to the folded state. Each
one of those pathways will have
different kinetics and different time
scales, but on average altogether they
make life possible. So it's robust
enough that on average a protein will be
folded every x number of units
>> because of that understanding how that
is actually the path to taking will make
us maybe think about drugs that prevent
miscaping. Right? Yeah. So, so I think
it's good to have the difference between
understanding and making predictions.
>> Yeah.
>> Thanks.
>> So, like
thank you very much for the lecture.
That's like what you come from the
simple
explanation to the complex ideas and
that's that's very nice. Uh I have a
question about like
you want to know how to come from the
structure to sequence and uh maybe
why do we need it right now? Maybe we
needed first to understand how the
structure calls the function and then
understand how to come from the
structure to signals,
>> right? It would it would be very so
like let's go back to the to the D,
right? Like say alpha fold for the DAG.
It would predict the D with its wings
stacked. That doesn't tell me anything
about what what is it that the D needs
in order to fly. Right? I need to
understand this other state which is the
flap the wings open and down and up. So
those are the three states I need to
understand for that right. If I
understand those three states and what
are the kinetics between them then I can
understand how it functions how it
flies.
>> Yes. So so understanding the u like the
path from the structure to sequence is
the key for understanding the structure
to function.
So one more time. So going from
structure to sequence is key to
>> understanding structure how the we can
define structure function function of
the partition.
>> Not necessarily I mean like I I think
like they are independent tools but all
together we will get more knowledge to
try to understand function. So what is
the degenerate code of sequences that
are able to fall to a particular
structure and what are the states that
they are compatible with and like how
does that relate to function that will
be interesting but you know I think it
all plays together but not necessarily
we need this to learn the other thing.
>> Yeah thank you very much. Yeah.
>> Any further questions?
>> Um, well, if not, yeah, thank you very
much again for
thank you for coming.