Bioexcel webinar #98: Conformational ensembles of intrinsically disordered regions and proteins
Watch on YouTubeVideo summary
Kerstin Lindorff Larsen from the University of Copenhagen introduced computational methods designed specifically for studying conformational ensembles within intrinsically disordered proteins (IDPs) and regions, emphasizing that while only about 5% of human proteins are fully disordered, most contain a mixture of ordered and disordered segments. Traditional modeling approaches relying on large structural databases like the PDB or multiple sequence alignments often fail to capture this biological context because IDPs lack well-defined native states and evolve differently than folded proteins. To overcome these limitations, researchers developed coarse-grained simulation models parameterized directly from experimental data such as NMR paramagnetic relaxation enhancement and small-angle X-ray scattering (SAXS), utilizing Bayesian optimization to fit force field parameters without depending on predefined ensembles or synthetic HP-models initially used in early iterations.
The resulting frameworks, including the Cabadas model which incorporates non-electrostatic interactions and hydrophobicity priors derived from literature scales, successfully predict chain compaction based directly on sequence features like charge distribution and blockiness. This capability allows for the rapid analysis of thousands to millions of sequences without extensive simulation time, effectively bridging the gap between folded and disordered states across a continuum defined by chemical shifts and SAXS data. A key application involves integrating these models with tools like AlphaFold2 or ColabFold; while standard AlphaFold excels at predicting specific structures for folded domains but struggles to generate the necessary structural diversity required for IDPs, newer approaches freeze well-folded regions using harmonic constraints while allowing disordered segments to be simulated dynamically via direct experimental data fitting.
Further analysis highlights that current models rely exclusively on solution-based experimental data rather than crystallographic information and balance electrostatic interactions with hydrophobicity through screened Coulomb potentials featuring fractional charges near pKa values, despite some limitations regarding salt dependency or varying dielectric constants. Although all-atom molecular dynamics offer detailed local structural insights, they often struggle with convergence for long IDPs exceeding 40 residues, making these efficient coarse-grained models superior for studying global properties and multi-domain proteins where disordered regions behave distinctly in full-length contexts compared to isolated loops. Future developments aim to expand benchmarks across mixed-order systems, integrate post-translational modifications and molecular crowding effects, and enhance nucleic acid modeling by adding harmonic constraints for duplex structures beyond current bead-based representations lacking sequence specificity.
The presentation concludes with a comprehensive evaluation of generative models like Pep-Zero against established benchmarks such as PeptoneBench, demonstrating that these new tools outperform traditional methods specifically on IDPs while maintaining accuracy on folded proteins when combined appropriately. By addressing the critical need for structural diversity and avoiding the loss functions inherent in standard AlphaFold approaches, this research provides a robust pathway to understand how intrinsically disordered regions function through chemistry-specific polymer properties and short linear motifs. The session also acknowledged the development team's contributions, invited further discussion via an online forum, and announced upcoming deadlines for the BioExcel conference events scheduled in Brno, underscoring the collaborative effort required to advance computational biology in this complex field.
Read the full video transcript
Welcome everybody.
Today we have the BioExcel webinar
number 98. And with us is Kerstin
Lindorff Larsen from the University of
Copenhagen.
And he will speak about conformational
ensembles of intrinsically disordered
region and protein.
I'm Alessandra Villa that together with
Otto Anderson and Richard Norma we are
hosting this
webinar for BioExcel Center of
Excellence for computational
biomolecular research.
This webinar is being recorded. I wanted
you are aware of that.
During the webinar you could ask
question using [snorts]
the Zoom application that you find at
the top bottom of your application.
Depends on the operating system you
might see one of those three symbols.
Please click it, type your question, and
at the end of the webinar
we will
read the question for you and Kerstin
will answer.
After the webinar maybe you still have
some question. You are most welcome to
join us at the BioExcel
ask.bioexcel.eu
forum.
And there will be a post
for the webinar of Kerstin and then you
can type your question there.
Kerstin will be able in the following
week to answer for the question.
Something about Kerstin. Kerstin is a
professor of computational protein
biophysics at the DeLund Strom Center
for protein science. He's director also
of the PRISM
protein interaction and stability in
medicine and genomics center.
And the member of the Royal Danish
Academy of Science and Letters.
He got a Danish Independent Research
Council Young Researcher in 20
in 2006. He was co-recipient in
2009 of the Golden Bell Prize.
Congratulations. And in 2025, the Novo
Nordisk Foundation Prize for natural
science teachers at the university.
Currently, his research focus on
developing applying computational
methods for studying studying the
structure and the dynamics of protein
and integration of biophysics and
genomics genomics research.
And today we'll speak about
intrinsically disordered region and
protein.
Okay.
I stop sharing and welcome.
>> Thank you very much, Alessandra, and
everyone else who's online for joining.
Uh, it's my pleasure to be here today to
talk about some of our work over the
years on developing
computational approaches to study
intrinsically disordered proteins.
So,
I
have put together
I let me just minimize this thing before
I
I've put together a series of slides
that talks about our work on how one can
determine conformational ensembles of
intrinsically disordered
proteins
and and regions and
focusing a little bit more on how we
develop the approaches
and then with with some pointers along
the way also on how one one one one can
use them.
Um so I'm not going to give a huge
amount of of biological background, but
nonetheless
just want to remind everyone that
proteins with disordered regions are
pervasive.
And so again, if I
can manage to get rid
of
things here.
Here's some some statistics that that
that other people have collected and and
you know, I've annotated it here with
sort of three uh sort of types of
proteins. Uh these are numbers from the
human proteome.
Um so if we sort of define proteins that
are sort of mostly folded as proteins
that have you know, at most I think 5%
of disorder, that turns out to be about
uh 40% of human proteins. On the other
hand, about 5% of human proteins
basically do not have uh well-defined
folded domains, which means that the
majority of proteins in the human
proteome contain mixtures of ordered and
disordered regions.
And so if the toolbox you have available
only works for fully folded proteins or
only works for fully disordered
proteins, you're basically losing out of
studying most of human proteins. And so
what I will talk about is how we are
building up computational models that
allow us to study this continuum of
protein order and disorder.
Another way of looking at the same
numbers is if you look at the amino acid
level, about 30% of the amino acid
residues in the human proteome are what
we would say uh intrinsically
disordered.
These proteins do a wide range of
things, and I put in this slide mostly
as an advertisement for this wonderful
review that Alex Holehouse and Peter
wrote a couple of years ago about
protein disorder ranging from the
biological function to the molecular
biophysics. And so, if you're interested
in these proteins, you should definitely
read this review as well as other nice
reviews from over the years. And of
course, the the the original literature
also. So, proteins with disordered
regions are involved in basically all
biological processes ranging from
cell signaling
to to transport
to you know, transcriptional regulation,
etc. etc.
And the way that they do this is via a
combination of different types of
chemistries. And some of these
chemistries are Well, these rules are
somewhat different from the rules that
we're used to think about if you think
about folded proteins. And so, one way I
think about it is that there are these
disordered regions that can either be at
the ends of proteins or between folded
domains that have chemistry-specific
polymer properties, which means that
it's not so much the specific amino acid
sequence, but it's more the overall
properties of these these regions that
drive some part of the function. And
then there are other parts of function,
for example, in so-called short linear
motifs that are more sequence-specific
that often govern more more specific
interactions. This is at the very high
level, but of course, in in in in many
individual cases, we also know exactly,
you know, how these different
properties, you know, play together and
how they determine biological function.
One of the difficulties in moving
forward in this field has been that the
rules and the techniques that are used
to sort of read and write sequences are
somewhat different than the rules that
we are sort of getting used to for for
reading and writing protein sequences
for folded domains. And so, I think it's
sometimes constructive to contrast the
situation of disordered proteins from
the situation for folded proteins.
And so if we for example think about the
data sources that have been been used to
develop structure prediction models or
protein design models for folded
proteins, they consist of two sources of
information. We have the wonderful
protein data bank with a few hundred
thousand examples of sequence structure
pairs, and then we have you know large
large sequence databases, UniProt, etc.
that have lots of sequences, and then
these sequences we can then align in
so-called multiple sequence alignments,
and then we can train deep learning
models that relate sequence alignments
to structure, either sort of to predict
structure or to generate sequences.
Or we can use other models that maybe
don't explicitly rely on alignments, but
often implicitly rely on the fact that
sequences are conserved in very special
ways.
For disordered proteins, we don't have
these. We have the protein ensemble
database that's small and heterogeneous.
We have lots and lots of sequences, but
it turns out that the way that these
develop over evolutionary time scales
are quite different from the constraints
that operate on folded regions, and so
they can
change the sequences in very different
ways, which means that multiple sequence
alignments sometimes give you useful
information, but other times the
alignments really do not tell you the
picture. And so we've been thinking
about not just our group, but many other
groups have been thinking about how can
we learn
sequence ensemble or sequence property
relationship for disordered proteins in
other ways, and this is what I'll talk
about today.
Some of these thoughts are summarized
in this recent review that we wrote, and
so this will be sort of the outline of
at least part of the presentation. And
the the major point is that if we don't
have huge amount of structural data, but
we still care about understanding
structural properties, how can we learn
those from the data that we have? And
so, what I'll talk about is how we can
take experimental data mostly from NMR
and small angle x-ray scattering data,
and use that to derive physical models,
in this case coarse-grained simulation
models that we will parameterize based
on experimental data, and I'll tell you
about how we do this.
And then how we can validate these
models that then also turn out to be
relatively fast.
And then we can use these to run
simulations and generate conformational
ensembles of thousands of of sequences
that we can then use as input to train
machine learning models if that's what
we want to do. Or we can just use the
simulations to give us some structural
insight. Or we can use them to interpret
experimental data. I will not really
talk about that, but but but I will put
in a a pointer at the very end.
We can also take these models and sort
of invert them either with the
simulation tools or with intermediate
machine learning models to do sequence
design, and again I'm not going to talk
about it, but it's described in this
review as well as another review we
recently wrote.
So, the goal is how can we learn when we
don't have a huge amount of experimental
data?
And I decided for this talk to go back a
few years sort of to explain you how we
got into this and some of the early work
that we did um and how this sort of
influenced our work later on. So, this
is some work that was done by a
fantastic master student I had in the
lab now about 20 years ago when we
started working on this problem. So,
previously I've been working on
disordered proteins and how we can
interpret experimental data on these,
but back at the time we thought maybe we
can develop prediction tools that we can
then sort of use to study when we don't
have experimental data. And so, Anders
together with a collaborator, Jasper
Borg, uh came up with a procedure of how
we can parameterize a coarse-grained
model for what we call structure
prediction of disordered proteins. Now,
this was done in the context of Monte
Carlo simulations, what I'll talk more
about later will be, you know, molecular
dynamics, but but but but the basically
the models are
are very similar. So, we want to have a
model that is, you know, as well as we
can, be predictive. And the work is
published again some years ago, but I'll
give you some few snippets just to to
give you, you know, an idea of of how we
we think about these problems and how
we've been sort of developing our
thinking.
So, if you have a model like this, you
need to parameterize it. And one of the
ways that we can parameterize
coarse-grained models is what we call
sort of bottom-up parameterization. We
have a more fine-grained model that we
can then represent. This we could do for
the backbone potential in a
coarse-grained model. And in in for
those of you who don't think about
C-alpha level models, the C-alpha based
models have a backbone that you can have
a backbone potential for. In this case,
there will be an angle and a torsion
that determines the local geometry.
We could actually generate reasonable
backbone structures using a series of
tools. Back then, there were these
statistical coil models that operate at
the full atom backbone level. And then
we can project those down into this
angle and the dihedral angle uh so, the
angle and the dihedral angle. And then
we can parameterize the coarse-grained
model using these mixture models here
that we can then combine to create what
is basically the equivalent of the
Ramachandran map, but for a C-alpha
chain.
So, this is a relatively simple way of
generating a a force field is to start
with a more fine-grained model and sort
of, you know, parameterize it
uh into into the coarse-grained model.
But the more important part of the model
is really the long-range or the
non-bonded parts. And we decided that
this would be better to do by tackling,
you know, some some some some some
experimental data. And again, these are
some old methods and you know, these
days we don't do these kind of things
anymore, but just to remind you that
there are a bunch of ways of creating
statistical models for force fields,
but very often they rely on having sort
of target structures. We might want to,
for example, you know, minimize, you
know, or come up with a force field that
you know, optimizes the Z-score so that
we can distinguish the native state from
other states. Or there's this fantastic
old work on funnel sculpting where you
maybe want to correlate the RMSD to the
native state with the energy. And you
can optimize the force fields against
these things. And there's a bunch of
different ways of optimizing, you know,
force fields that, you know, stabilize a
well-defined native state relative to
some other states. But all of these
models generally relied on some
well-defined target structures. And of
course, for these folded proteins, we
don't have a native state, or rather
native state is a broad conformational
ensemble. So, we thought that these
models were not really the right way of
thinking about it.
We and many others have been thinking
about how one we can generate
conformational ensembles. I'm here
showing you two conformational ensembles
of a protein. It's not so important what
what it is. What is important is that
these are two conformational ensembles,
one on the left and one on the right,
and they're both derived from the same
experimental data.
In this case, from NMR experiments. And
you can see that they're quite
different. So, it turns out that there's
actually quite many different ways of
generating conformational ensembles that
are compatible with a single set of
experimental data. And so, we were not
super comfortable on targeting the
experimental ensembles. So, we could
generate experimental ensembles based on
the data,
but given this uncertainty of the
ensembles, we thought, well, maybe this
is not the right approach.
So, instead, what we came up with was an
approach of optimizing the force fields
directly against the experimental data.
And so, in this paper that that Anders
published, uh really maybe the most
important part of the paper is is
described in the supplementary
information, and it's a Bayesian
approach where you formulate the
learning of the force field parameters
as an optimization problem given some
set of experimental data.
And so, the equation here at the top is
basically base formula that says, what
is the probability of some force field
parameters epsilon given the data can be
formulated as you know, relating to the
probability of serving some calculated
data or some data given the force field
parameters.
And this is something that we can then
calculate because if we have some guess
of what
what the force field would look like.
So, if you have some hypothesis of what
the force field would look like, we have
some hypothesis of what these epsilons
would be, we can then run simulations,
and then we can calculate whatever we
might observe in an experiment, and then
we can compare that to the experiment
that we have. And so, this is basically
what we derived. The equation below is
the same as the above under some
simplifying assumptions and taking some
logarithms. And basically, the idea is
that we have sort of a chi-squared term
that measures the deviation between what
we measure experimentally and what we
predict from a simulation, but where the
prediction from the simulation is
dependent on the force field.
And so, we can optimize this
log-likelihood function basically by
minimizing this. So, that's the basic
idea. Of course, it's it's it's it's uh
conceptually uh simple to write down.
The difficulty is how you implement this
in practice. And so, this is what what
Anders worked on.
Before I get to that, I'll just, you
know, showcase you one of the types of
the NMR experiments that we can do. This
is paramagnetic relaxation enhancement
NMR data. This will just be important
because that's the few points of data
that I'll show you. If we put in a spin
label at a site, we can see broadening
of NMRs throughout the chain in a
distance-dependent fashion. And again,
under some simplifying assumptions,
there's a simple relationship between
the average distance between the spin
label and the protons or the
distributions of distances that gives
you this paramagnetic relaxation
enhancement term that then influences
these insensitive ratios. And although
this is not the perfect measure, this is
what we had available when we did this
work again a long time ago. So, if we
have a conformational ensemble, we can
calculate the PIE and we can compare it
to experiment. So, this is the
experiment that I'll I'll I'll talk
about.
And so, we came up with this iterative
optimization scheme where we have some
experimental data. I'll show you how it
works with the synthetic data. And then,
I'll show you how we can optimize
against this and show that the method
works.
And so, the idea is that we start with
some parameters that we, you know, that
we, um, that we guessed. They could, for
example, see we set all the attractions
to zero.
We run a simulation to generate a
conformational ensemble with that
initial parameter set.
We calculate the experimental data as,
you know, observables.
And then, the key trick is that then we
can optimize the parameters
so that we minimize the agreement or
maximize the agreement or minimize the
difference between what we predict from
the simulation and what we measure
experimentally and use that as a guiding
function for the force field parameters.
And then, we can iterate over and over
again.
And then, the whole trick is how do you
do this without having to run so many
simulations to test all possible force
fields. That's really the difficult
part.
So, again, to test that the model work,
honest came up with a nice synthetic
data set. He said, "Well, the simplest
way we can represent a protein is sort
of what we might now call a stickers and
spaces model." Or in the protein folding
field, uh for for polymers, this was the
the HP model. Um but it's basically a
model where there's two types of
residues. There's hydrophobic residues
that interact favorably with each
others, and then there's polar residues
that don't interact favorably with each
other or with the hydrophobic residues.
And so, we created a synthetic uh data
set. These are the black lines here that
if we took this protein called ACBP and
simulated it, we would measure this data
set. So, this is now our synthetic data
set. And then we say, "Well, we don't
really know this one, so we'll start by
setting all the interactions to zero."
And the red lines is what you get if you
run a finite-length simulation, you
know, so with some noise um using a
model where all the interactions are
zero. And so, we ask the question, "Can
we recover this HP model um from this
experimental data?"
And so, the approach is again relatively
similar. We run a simple We run a
simulation. We calculate the expe-
experimental observables as these
weighted uh averages here. Now, these
weights here would typically just be
uniform weights because we're doing in
this case Monte Carlo
uh sampling. And so, we can just run
simulation. We can calculate these
numbers. We can compare with the
experiment, so we can calculate a
chi-squared term
um for what these these these these
would look like.
And now, the trick here is that we can
now actually calculate the derivative of
this chi-squared term with respect to
the force field parameters. In this
model, the force field parameters are
called Q. So, we can calculate the
derivative of this with respect to the
force field parameters to do a gradient
optimization
where we minimize this chi-squared term
by optimizing these Qs.
And of course, this is difficult because
we have a big conformational ensemble,
so we need to propagate the changes in
the force field through the
conformational ensemble, but then
Boltzmann comes to our rescue, so we can
basically, at least for local changes,
we can calculate how the energy of each
of the conformations would change if we
update the force field. So, we can
propagate the
changes in the Qs back through the
ensemble to optimize this. And so, the
gradients back then, we would calculate
using finite differences. Then, of
course, these days one can do this with
automated differentiation tools, uh but
back then we did it, you know, using
this simple approach. And then you can
then basically optimize the force field
at least locally, and then you have to
resample using the optimized parameters,
and then you iterate over and over
again.
So, how does that look in practice?
Here's one example.
Here, the black line is as before, this
is the synthetic experimental data. The
green line is what you get from a force
field that is bad.
But you can then take this green line,
and you can re-weight it by optimizing
these force fields,
and you can do that at least to some
extent, and you get this noisy red line,
where we optimize the force field
parameters to get closer and closer to
the experiments without doing any
resampling. This gets noisy because we
are stuck with the original ensemble,
but once we are close enough, and that's
described in the paper what close enough
means, we can then rerun a simulation,
we get this blue, and you can see this
blue line is now closer to the
experiment than the green line, meaning
that we optimized the force field in one
single step.
And then we can basically do this over
and over again. This is what's shown
here on the right. We start here with
the red, and then we move continuously
closer and closer to the black line,
which is the experiment. And we can see
over here what happens with the force
field. We start off with all the force
field parameters being zero, and very
quickly this optimization model learns
that this HH interaction should be
favorable with a minus one energy, and
the other one should stay around zero.
So, this method sort of works in the
sense that we can recover a synthetic
force field that we predefined.
This is a simple model with, you know,
three parameters. We also showed that we
can recover a 20-parameter model, and we
also showed how one can use it for
experimental data.
So, this is basically a approach of
learning force field parameters directly
from experimental data without having to
define conformational ensembles.
And so, one might have think, well, that
was the, you know, the end of it, but of
course, for us this was really the end
of it, but only because we sort of
didn't really continue this work. And
one of the reasons was that at the time
there was really not very much of this
kind of experimental data available.
And while there was some experimental
available, this was often done on a very
different set of conditions. Lots of the
data were for unfolded proteins that
were denatured in denaturants and acid,
and there was not a lot of
systematically collected data for
intrinsically disordered proteins, at
least not that we were aware of.
Um but over the years, this has
certainly changed. And so, over the
years, lots of other people started to
do similar things, you know, to develop
coarse-grained models either from
top-down or bottom-up, and validate
these using experimental data. And then,
some years ago, we decided to, well,
maybe we can revive uh
this idea of learning force fields
directly from experimental data, but
using newer tools, faster models, and
most importantly, a larger set of
experimental data.
One of the reasons why we and others
also uh got back into this question was
that there was a research interest in
studying just not just individual
disordered proteins, but also
interactions between disordered
proteins. And so this figure is sort of
meant to illustrate you that, you know,
if we learn the rules of intramolecular
interactions, so between different amino
acids within the single chain, we also
learn something about the intermolecular
interactions between chains. And this
turns out to be important if you want to
study cluster formation or for example
phase separation. And although I'll not
talk about this, I'll just show you here
some some some some experimental and
computational data from a bunch of
groups. These are sort of some key
papers, but there are several others
from over the years and also prior work
in the in the polymer field. And then
each of these plots, the stuff on the
x-axis is related to how compact a
disordered protein is, either by
measured experimentally or from
simulations of various types or
competing or theoretical models. And on
the y-axis are various things related to
the protein's tendency to self-associate
and for example undergo phase
separation. And these correlations show
that there is this intrinsic
correlation. And so we use that as as
sort of motivation to say that if we can
learn a force field by looking at single
chain properties, so this is the stuff
on the x-axis,
maybe it tells us also something about
these proteins' tendency to undergo
phase separation. And there's of course
other people had had thought about this
and sort of we just piggyback on on on
these ideas.
And so the idea was sort of picked up
again initially by Ramon and Tea and
then later joined by several other
people, most notably Julio, and but
Fanny, Key, Soren, Ariane, Anaida and
and and several others have contributed
over the years to sort of
redeveloping this data-driven
coarse-grained modeling. And so the
model I'll talk about briefly is called
Cabadas. It's a copy of other people's
models. It has some similarities to our
our previous models, but also some some
additions. So for example, there's now
molecular dynamics, there's now charge
interactions, and and a few other
differences.
Um but but the key key parameter in this
is still these non non-local
interactions between amino acids, in
particular these non-electrostatic
interactions that are parameterized by
this parameter that we we call lambda
that says something about whether the
amino acids like to interact with each
other or whether they more like to
interact with their solvent that we're
not representing explicitly. And so we
call this sometimes a hydrophobicity
parameter or a stickiness parameter.
So again, we took out the the the the
model from from from Anders, we we
changed it in a few different ways, and
actually in some sense simplified the
parameter learning
um in a way that turned out to be a
little bit faster, but the basic idea is
the same. We write down this
log-likelihood function that depends on
the lambdas. We have two terms here that
measure the deviation to experiments,
and then one term over here that we
ignored in the previous work. This is
the prior in the Bayesian
formulation, and then we basically
optimize or minimize this function that
sort of, you know,
maximize the agreement with experiments,
and, you know, minimizes the deviation
from the prior. And again, it's this
kind of iterative, you know, sampling
and and and re-weighting kind of
approach.
So for the experimental data, we went
into the literature, we found
experimental data for initially about 50
proteins. Now we are in more than 100
proteins, mostly from small angle X-ray
scattering data measured by wonderful
collaborators and non-collaborators over
the world who measures data, puts them
in databases, make them available so
that we and others can can use it. So
thank you very much to all the people
that measure measure data. That's
fantastic. As well as NMR data again
that we've collected from the literature
or people have, you know, dug out of old
folders and and so on and so forth. So
we have a a medium-size data set that
allows us to learn these parameters.
Now, this updated version we also have a
prior and so our prior was to go into
the literature and look at
hydrophobicity scales. And so we we took
about 50 or 70 hydrophobicity scales,
you know, normalized them and say, well,
you know, maybe they don't agree all of
them, but most of them think that
aspartate is not so sticky and
tryptophan is more sticky. So, at least
in the absence of experimental data,
this would be our prior expectation. And
so we used that sort of to regularize
the optimization also. And so this in
particular plays a role for for for
amino acid
residues that are are less common across
these 100 proteins or so that that we
are now optimizing on.
So, the optimization looks a little bit,
you know, look pretty similar to what I
showed you previously. I'll just play
you a movie that sort of shows how it
operates. We I know, we run a simulation
with a force field and we're slowly
optimizing it so that we maximize the
agreement in this case, for example,
with the radius of observation or over
here the NMR data. Um but now doing it
not just on a single protein, but you
know, simultaneously on in this case
about 50 protein and later on, you know,
about 100 different proteins in in
total.
And so in total, we can optimize force
fields by
maximizing agreement with experimental
data on a broad range of proteins as
well as this prior that comes from the
hydrophobicity scales. And so this is
sort of roughly what the final
stickiness scale looks like and it has
some similarities and some differences
to the original model. We call the model
Cavitas, this is the acronym, it's not
so important. We we've we validated it
quite broadly as I'll show you in a
second.
This is sort of the agreement with the
training data set. You know, this has a
very high correlation mostly because
it's pretty easy to predict that long
proteins have a greater RG than small
proteins. But also if you look, for
example, of variants of the low
complexity of A1
from from mostly 10 nanometers lab, we
can see a very high correlation
between the measured and the calculated
radius of gyration. This is now within
the training data. So at least this
tells us that we can recapitulate these
things.
But we can also look at independent data
that other people have have collected or
we have collected over the years. Sex
data, fret data or or sex data
you know on proteins that all have the
same length
but where maybe the amino acids have
been moved around. And again this is
described in published literature. I'm
not going to go through it. Just to say
we collect data from the literature and
again you know big acknowledgement to
all the people who measure high quality
data and make it available.
And then we can use this in this case to
benchmark the models.
When we then look at this scale, we can
see that this scale sort of makes sense
in the sense that it looks like what
people had suggested by looking at
experimental data for disordered
proteins that there are some residue
types that like to interact more with
each other than others.
We've just published a paper where we
sort of try to make this systematic also
by
saying that the radius you know trades
we get get get rid of the problem that
the different amino acids have slightly
different radii.
But we can then also go back and look at
what kind of hydrophobicity scales to
these to this scale correlate with. And
what we for example we find that this
you know scales that are down here,
these are are sort of our new scales,
they are similar to some experimental
scales for example this Jure scale that
others have been using to study IDPs in
simulation that is parametrized again on
sort of disordered proteins but is for
example quite different from the
commonly used Kyte-Doolittle scale and
because some of the residues that are
sort of sticky or or or interaction
prone in the context of disordered
proteins are not interaction prone in
the context of of IDPs. And so, this
scale, you know, in many ways
recapitulate what other people had found
in sort of individual cases, but maybe
makes it also systematic across all
different 20 amino acid types, learning
it directly from the experimental data.
So, what can we do with such a model?
I'll show you uh you know, a few few
snippets of what one can do. So, we can,
for example, now start to explore
sequence-structure relationships. So,
how do we go from sequence to structural
properties?
In one example, Julio and Anna ran
simulations of about 28,000 IDRs that we
extracted from the human proteome,
cutting out the IDRs of of of
full-length proteins, and ran
simulations of all of them, put them in
a database that you can analyze. So, I'm
not going to go through the details.
This is all published. Just to say that
we can sort of conceptually group them
into things that have different levels
of compaction independently of chain
length. You can quantify this in
different ways. One way is to take this
apparent internal distance scaling
exponent, that when it's low, it means
that the chain is compact, and when it
when it's high, it means that the chain
is more expanded, but that is more
independent of the number of amino
acids. And we can then start to say,
what are the sequence properties that
determine this? And we find that these
again recapitulate what others had found
in individual cases, and sometimes also
systematically.
So, that there is this interplay between
having a high or low fraction of charged
residues, a high or low hydrophobicity,
and if you have many charged residues,
whether they are well-mixed or blocky
along the sequence.
And so, we could put all of this
together in a single coherent framework
that sort of combines these features.
Again, not features that we develop, but
put them into a single framework that
say, what is the interplay between these
across proteins that are natural
sequences that determine whether they
are highly expanded or highly compact.
And basically rules are pretty simple.
If you are highly charged and not
charged blocky, you are typically
expanded. If you are medium, you know,
if you hydrophobic and not so charged,
you are typically more compact. But if
you are highly charged, but then charged
blocky and a little bit less
hydrophobic, you can then become even
more compact by forming long-range
interactions that are driven by
charge-charge
interactions in these kind of more
blocky sequences.
These rules are actually relatively
simple, so you can develop very, very,
very simple machine learning model to
predict this number from sequence. This
is support vector regression, so this is
probably some of the simplest models
that you could develop, but it turns out
to work, you know, extremely well just
for predicting conformational properties
such as chain compaction directly from
sequence. So, we don't even need to run
simulations. We can just predict from
sequence, and this is available for
anyone to predict.
This is convenient because that means
that we can scale up instead of looking
at 20,000, we can look at a million
sequences, and we can ask questions. For
example, are these properties conserved
across evolution? And it turns out that
the compaction of an ortholog is
typically strongly correlated to the
compaction of a human IDR, not always,
but at least on average.
So, this, you know, was is maybe not
surprising. And of course, it depends a
little bit on how you find orthologs.
That's a whole separate talk. Um but it
turns out that others had found that
while this is true often, it's not
always true. So, here's a really nice
paper. This is not our work, that showed
that there was this intrinsic
correlation between the sort of per
amino acid uh compaction or in this case
specifically actually the end-to-end
distance and and the length so that
there is some
compensatory changes in the sequences
over evolution so that if the IDR
becomes very long, it becomes a little
bit more compact per residue. And so we
can now start to see can we find
equivalent cases like this
where it appears that the global chain
dimensions are more conserved than for
example the length of the amino acid
chain.
And so again, this is published so I'm
not going to go through it just to say
that there are several cases where it
appears that over evolutionary
timescales there is a correlated change
between the length of an IDR and sort of
the sequence properties that drive
compaction so that there's this negative
correlation so as the chain gets longer
across evolution or shorter for that
matter, then the sequences adopt to
preserve the physical dimensions for
example the radius of gyration or the
end-to-end distance more than for
example the average amino acid
composition suggesting that at least in
some of these cases the physical
dimensions might be under evolutionary
pressure like in the paper that I just
showed you previously where they
demonstrated you know this has strong
functional roles.
So I started by saying that only 5% of
human proteins are fully disordered so
of course it's important to put this
disorder back into context and we've
been pushing the model to work for
multi-domain proteins, nucleic acids,
we're interested in crowding, PTMs and
other things.
And I won't have time to talk about all
of this but I'll just show you briefly
some work we've been doing on
multi-domain proteins. We developed a
model called Calvados 3 where we keep
the folded domains rigid by a harmonic
network constraint how you might do it
in other coarse-grained models also. The
model predicts radius of gyration and
also phase separation relatively
accurately, so it doesn't sample the
dynamics of these folded domains
particularly accurately, but it keeps
the structure together as you would, you
know, in in in other coarse-grained
models.
One limitation of this is that we need
manually to define where the folded
domains are, and this can be a little
bit tedious if you want to scale up. So,
obviously we came up with an alternative
approach of doing this that we call
AlphaFold ColabFold, there's a preprint
out um that you can have a look at. And
basically, we replaced this manual step,
or rather that CERN developed this
manual step, where he basically combined
AlphaFold structure prediction together
with PLDDT score and the AlphaFold PAE
matrices to come up with a way of
automatically constraining the folded
domains and letting the disordered
regions be flexible and governed purely
by ColabFold. Then the folded domains
are kept together with sort of a go-type
structure-based model. And so, we can
now easily take a sequence,
run an AlphaFold prediction, and then
input that into the AlphaFold ColabFold
framework, and now generate
conformational ensembles of a protein
with both folded regions and disordered
regions without having to do any manual
assignment of the folded state.
This model is as accurate as the manual
model, so that's great, but it's
automated so that CERN could easily run
simulations of in this case 12,000
full-length human proteins that are, you
know, found in the cytoplasm or the
nucleus. And so, we put them in a
database that you can go analyze.
Just to give you one example of the kind
of analysis we can do, we can now
compare how the disordered regions
behave when the context of the
full-length protein versus in the
context of, you know, just being
isolated. And we're sort of mining this
data to try to figure out what are the
context rules. For example, we find some
ideas that are much more compact in the
context of the full-length protein than
in when they're isolated. These are not
big discoveries. This is what we
commonly call loops. So, these are
basically regions that are constrained
because they're sitting either within a
folded domain or between folded domains
that are tied together. So, this is not
a big discovery, but they fall out very
naturally from this kind of analysis.
But, there's lots of other variation
that that is interesting and hopefully,
you know, we and others will be using
our data. This is our data for
transcription factors.
And in the final few minutes, I'll just
say that, you know, for these kind of
models and you other models out there,
we really need better and broader
benchmarks that operate not just on the
disordered regions that I showed you we
have the SAXS and the NMR data, or the
folded regions where of course we have,
you know, the wonderful PDB, but also of
things in between.
And so, a part of this other paper that
developed a generative model either for
IDPs or for folding protein, this is
work driven mostly by people at the
company Peptone together with NVIDIA,
and we also played a minor role in in
this work. So, I'm not going to talk
about this model, but as part of this
work, one of the key steps in this paper
was also to put together a benchmark
that operates across this order-disorder
continuum.
The model is called PeptoneBench. It has
several hundred proteins that have been,
you know, mined from the two databases,
the chemical shift database BMRB and the
SAXS database SASDB,
automatically extracted so that we can
now take structure models for proteins
with mixed order and disorder and
benchmark both their local structure by
comparing to chemical shift and global
structure by comparing, for example, to
small-angle X-ray scattering data. So,
we can now calculate chemical the
and compare to experiments, and then put
it on this RMS E axis, so when low means
good and high means bad. But now, each
of these points is an individual protein
or individual data set that sort of
spans this order to disorder continuum.
And this means now that we have
generative models or simulation models,
we can now benchmark these at scale.
So, here's just some example what this
might look like. The top row is for
chemical shift, the bottom row is for
small angle X-ray scattering data. If
you take AlphaFold just as a simple
baseline for structure prediction, a
folded protein, AlphaFold, not
surprisingly, does amazingly well for
folded proteins, but of course doesn't
really predict chemical shift of
disordered proteins because that's not
what it was developed to do. That's true
also for SAXS data.
If you take this IDPO model, that's IDP
specific, it predicts SAXS data very
well for disordered proteins, but not so
well for ordered proteins.
Whereas this Pep-Zero model, for
example, operates relatively well across
the order-disorder continuum.
So, we can quantify how well something
works by basically calculating the area
under this curve from the ordered
regions to the disordered regions. And
here's some benchmarks on different
models. There's AlphaFold and both,
there's BioLuminate Pep-Zero, and you
can see, for example, over here for
chemical shift over here for SAXS, that
these more broadly applicable generative
models do quite a bit better across the
entire order-disorder continuum. So,
whereas AlphaFold does fantastically for
folded proteins because that's what it's
meant to do, it doesn't do so well over
here. And so, we said, "How well does
our model do on this?" And we were very
pleased to see that by basic
copying structure of the folded domains
from AlphaFold
in AlphaFold ColabFold, but then letting
ColabFold deal with all of the disorder,
we actually created a model that for
global properties, as measured by SAXS,
actually does really well across the
entire disorder continuum. So, it even
does better than this bio immune and
peptide model and these in this kind of
benchmarks. So, this model, which is
very simple, it has 20 parameters,
um basically that interaction strength
for each amino acid type, allows us to
predict conformational ensembles of
proteins with mixed order and disorder
that is at least as accurate as these
more complicated models. Of course,
this model doesn't do everything that
these models do, but it's very efficient
and and and and and and and and as you
can see, pretty accurate at capturing
these conformational changes.
You can run all of this by going to our
GitHub page, either download the
Calvados package and there's sort of a a
manual here or a bunch of examples. You
can also run it through Colab. Um
there's the AlphaFold Calvados where we
put in all the conformational ensembles
that you can just download. If you know
the UniProt ID, you can click through
and generate our simulations. Or down
here at the end, you can also run it
yourself with a Colab certificate.
And so, with that, I'm at the end.
We can learn force field parameters uh
by targeting experimental data.
Um
we can combine this with traditional
machine learning models such as, for
example, AlphaFold protein language
model to build in specific interactions
and we can run these simulations at
scale and then we can generate training
data for predicting properties from
sequence. I showed you one example of
this. I didn't talk about how we can use
it for generative models. Um or we can
invert these models again for sequence
design. And again, I just want to say
that we need better and broader
benchmarks that are based on
experimental data. I showed you one data
set. I think it's very good, but we need
more than this. I I think I mentioned
the people along the way and that did
most of this work, um but I'll just show
you their faces here. I'll show their
faces here at the end um for the people
who developed first the early version of
Calvados and now the more recent
versions, as well as the people that did
all the applications that I talked
about. With that, I'll stop and happily
answer any questions.
>> Thank you very much, Kerstin. It was a
great presentation. We We got also a lot
of question in the chapter, yeah.
So, I will So, some of them are tied to
a
exactly moment in your presentation. So,
I don't know if they are understandable,
but I will try to
to be to go to jump from one to the
other.
And so, one at the beginning, people
were wondering in general, why do you
you What do you use
for experimental data?
Because you probably can't use any
crystal data or this for disorder
protein.
Yeah.
>> Yeah, so the
we have basically only used solution
experimental data, and that would mean
in almost, you know, close to 100%
either small angle X-ray scattering or
NMR spectroscopy data. For the
scattering data, if people all deposited
the data, we could use the scattering
curves. Often, we have to resort to
using the radius of gyration. For NMR,
we mostly use paramagnetic relaxation
enhancement, although we've also done
some work using chemical shifts.
We've done some work, not that I showed
you on optimizing against single
molecule fluorescent experiment data.
But basically, always we use solution
data that reports on the conformational
properties, at least in the stuff that I
presented here here today.
>> Oh, thank you. So, another question is
What would you consider the main
limitation of AlphaFold like approach
when applied to intrinsically disordered
protein? This is in the disordered
regions, sorry.
And there is a following up question in
the contest still of the intrinsic
disordered region, do you generally find
apple or holo system to be the more
informative for understanding their
functional behavior?
>> Yeah, so if you take the first question
first, I I think that there is, you
know, multiple layers to this. I I think
that, you know, if we take
AlphaFold-like models, it's a very broad
term.
I think the main limitation that, you
know, in many of the applications is
that they are, you know, they are
trained with a loss function against a
specific structure. And that
can cause all sorts of problems if you
have something that has a conformational
ensemble. And so what I tried to explain
is that, at least the way that we've
done it is that we find it more relevant
to target directly the experimental
data.
And of course, there are some approaches
where you can optimize models such as
AlphaFold, not by targeting XYZ
coordinates, but by propagating all the
way back to the, for example, the
crystal data, but this is not so common
anymore.
And then of course, the models
themselves need somehow to generate
diversity. And you know, there are now
lots of deep learning models that
actually can generate diversity, but
that has to be explicit in there also.
So you need to have a model that
generates diversity, and then your loss
function somehow needs to, you know,
include this diversity when comparing to
whatever training data you have.
As for the kinds of systems, you know,
for most of the things we've looked at
here, we really looked at at things that
are very dynamic. So we've looked at the
most disordered parts of it, and so we
don't really,
you know, we don't really focus so much
on whether the folded domains are in an
apo or holo state. And so for example,
for everything I've showed you with
folded domains, we basically freeze the
folded domain in whatever structure we
think is most relevant for the
conditions under which the experimental
data has been generated. So, if the
small angle x-ray scattering data is
generated for the apo protein, we
simulate the apo protein. If it's
generated with the holo protein, we will
try to simulate the holo protein. But,
we keep the folded domains more or less
rigid.
>> Okay. Thank you very much.
We have
Sorry. I just said
Uh we we have another question. I
thought we have a lot of questions, but
I just pick up
another one that is considering
intrinsically disordered protein have
more charged residue,
how use of a model predominant on
hydrophobic interaction can be
justified?
>> Yes. I mean, that's a a great point. So,
in the
first model that I showed you,
there was no electrostatics, and that
was clearly
you know, suboptimal. In the model that
we now have, the Calvados model, there
is electrostatics. So, charged residues
do exist. We treat them very simply. So,
we don't have
you know, you know, residues have a
charge, and so we don't represent the
charge distribution. If they are close
to their pKa, we can give them a
fractional charge. So, there's already a
series of approximations that we put in
there.
And then these charge-charge
interactions are are represented with
sort of a screened Coulomb potential
with a dielectric constant that we set
to be a constant depending on the, you
know, dielectric constant of solvent
with a Debye-Hückel term. And so, we do
have electrostatics in there, but it's
not very refined.
And I think
there are many cases where this will
fail. Um
I will say I've been relatively
surprised in how well it works and this
is not just what we found, other people
have found this also. So, I showed you a
couple of of examples of both other
people's data and our own data where we
can take a sequence and we can move
around the amino acids including moving
around the charged amino acids to change
the compaction and the model actually
captures this very well.
It doesn't capture particularly well
salt dependency of compaction. It does
to some extent. So, there are clearly
some limitations in how we treat
electrostatics and of course, if you
have environments that have different
effective dielectric constants, it will
also fail.
But, we do include it to, you know, to a
level that I think is comparable in
resolution to
other parts of the model and I think we
balance the electrostatics and the sort
of hydrophobicity terms relatively
reasonably in the model.
>> Yes, thank you. I have a We have a There
are
one other question one question on the
multi-domain model.
Uh,
do you do you put very strong harmonic
restraint in your folded domain or does
your model also allow to some
fluctuation conformational change within
the folded region?
>> Um, the answer to to both questions is
yes.
You know, so in the Calpha 3 model, we
have harmonic constraints in the
AlphaFold Calpha 3 model, we have to use
a more like, you know, Lennard-Jones
like a 12-10 potential as, you know, is
more commonly used in in in in, you
know, what originally was called Go
models or more general class of
structure-based models. Um, I will say
that we typically set them to be
relatively strong um,
to to sure that we keep the folded
domains constant, we optimize this
against the experimental data.
You know,
some years ago we published a paper
where we combined go models with other
physical force fields, sort of try to
you know, model folding unfolding
equilibria. And in principle, we could
do this here also.
Um, but in this case we said we we would
rather err a little bit on the on the
side of you know, keeping things
restrained. So, we do model some
dynamics, but I will say that the way
that we model the dynamics is mostly
that if we have loops, they are more
flexible than secondary structure
elements, and we don't try to optimize,
you know, local fluctuations
particularly accurately. If you are
interested in some specific system,
there is quite a bit of room of of of of
treating these parameters, but for
single parameter that works across all
proteins, we went for something that is
relatively conservative.
>> Okay, thank you.
And there is another question, um, that
is
uh, do you see any bias in the
persistent length of conservative
intrinsic disordered region versus not
conservative linker
pro linker versus protein
that are full
intrinsic disordered region that has
full disint- intrinsic disordered
regions, I think.
>> I
I think we've not really analyzed this
in any way, so I will just say again,
everything uh, that I showed is publicly
available. All the thousands of
simulations are available for anyone to
analyze.
I will say that I while I started by
saying that 20 years ago we had a local
potential that I think we optimized
relatively well in the current model, we
actually don't. And even that original
local model was not sequence dependent.
So, I'm not sure
how well we actually model sequence
variation in persistent length beyond
what is governed by these hydrophobic
interaction parameters? So, there is and
other people have done this. There is
room for improvement on developing
models to capture this that can then
potentially address the question that's
being asked. So, it's a great question.
We've not looked at it, but I would also
not be sure whether the current model
has the accuracy that could really
address this.
>> Yeah. Thank you very much.
And how do you compare atomistic
molecular dynamics with Calvados
constrained molecular dynamics, I guess,
for monomeric study of protein?
>> They have
different use cases. We still do some
all-atom MD of disordered proteins. They
are very, very difficult to converge.
When you go beyond, let's say, 30, 40,
50 residues, you really need to spend
quite a bit of computational resources
and and, you know, use non-trivial
sampling methods to converge them.
But, of course, they give you huge
amount of detail that our coarse-grained
models and other coarse-grained models
don't give us.
And so, for example, local structural
properties, coupling between local and
structural properties, detailed
interactions between side chains,
um all of this you get from all-atom MD.
I don't think that all-atom MD as it is
now is actually more accurate than the
coarse-grained models for measuring
global structure. So, I think that if we
could converge all-atom MD for these
very long IDPs that we're simulating and
we unfortunately can't, I don't think
they would agree better with
experiments, but of course, they agree
much better with experiments on local
properties and all um
and or, you know, giving you the
chemical detail that is often necessary
for studying interactions with small
molecules or or other proteins. There
are of course intermediate level models.
There's Martini-like models that have
some reduced
coarse-graining compared to, you know,
it's more coarse-grained than all-atom,
but less coarse-grained than Calvados.
Or there is sort of all-atom
implicit solvent models. And there's a
bunch of models for IDPs. There's
Absinthe, we developed one called EF1SB.
And there are several models that can do
these things that, you know, sit in this
space between.
They do different things, and I think
it's a balance of, you know, how much
computational resources you want to
apply, how broad you want to go in
sequence space, versus how much do you
care about a specific system.
>> Thank you. I reformulated the last
question that I will do. It's What's the
question was concerning nucleic acid, so
DNA or RNA
in the Calvados framework.
>> Yes. So, we have a published model of
what we'll call disordered
RNA. And we call it disordered RNA
because it has no secondary structure.
Each,
you know, nucleotide is represented by
two beads, sort of a backbone bead and a
as a base bead. There is no sequence, so
it's really a flexible polymer that we
optimize sort of the local potential to
match all-atom MD.
But of course, given that there's no
sequence, it's relatively rough. And
then the non-bonded terms are also
optimized to give us global
conformational properties. It's very
simple, so there's a huge amount of
things that it cannot do. But for some
things, it actually does reasonably
well.
We then also have unpublished work that
we hopefully will will be able to share
soon, where we can do similar to what
we've done for the Calvados 3 and alpha
Calvados model, where we put on harmonic
constraints to keep secondary structure
either in DNA or in RNA
um to study new systems where where
where you know duplex RNA or DNA would
be important. Again, these models are
currently not sequence specific. They
keep the structure rigid. We tune the
parameters to get the persistent length
and other interactions roughly correct,
but they are very primitive and I would
also say that they are you know much
more primitive than other models out
there um and they're more primitive than
the protein part of Cavas, but they may
be good enough to address some
questions. So, if you want to address
questions related to average stickiness
or charges or some kind of heterotypic
phase separation, maybe it's good
enough. And you know, we'll we'll keep
pushing these models, but I think it's
also important to remember that there
are limits to all models, be it all-atom
MD or different levels of coarse-grained
models, and there are just some things
that this kind of coarse-grained will
never be able to capture. And so,
there's a limit to how much, you know,
chemical detail that we can reasonably
try to put into a model that's as simple
as the models that I've been presenting.
So, we we have models for nucleic acids.
They are okay for some things and
terrible for other things. And
hopefully, we try to write in our papers
what we think that they can do and what
they cannot do. And if you're not
pleased, then let us know. We'll try to
do better.
>> Thank you very much, Kerstin. Thank you.
And there are still a lot of question. I
invite all the attendees that ask
question to use the Ask by Excel forum
that I put in the presentation to ask
the the question that you have here to
to type your question there, so Kerstin
can address those question. They will
answer question in the next week
in the forum.
So, I thank Thank for attending. I thank
you
Kerstin. It a fantastic talk. You got a
lot of nice talk mentioned from that
attendees. They were all enthusiastic.
So, thank you very much for join us for
this BioExcel webinar. Thank you to for
all attendees for being with us and for
the question. I thank you Autumn Richard
for join also this section and I will
remind all of you that BioExcel has a
BioExcel conference coming in September.
Please register. The deadline end of May
is the deadline for the early birth
and presentation of abstract. So, please
join it. You find all the details on
BioExcel webpage. I can I can also share
now the Maybe it's easy if I share
my presentation so you can see
the detail. Just give me a moment.
Uh
Sorry.
So, that is uh
So, that is the details. So, the barcode
is there. So, you can join us. The 31st
of May is the deadline for the early
birth. And please join us in Brno in
September
this year. And then I wish all the other
we stop here the spring section of the
BioExcel webinar and we will have a new
section coming up in autumn.
So,
thank you very much everybody.
Bye-bye.