Bioexcel webinar #97: Driving forces in biomolecular condensates from atomistic simulations
Watch on YouTubeVideo summary
David Sancho from the University of the Basque Country presented a reductionist approach to understanding biomolecular condensates by utilizing atomistic simulations on model peptides rather than relying solely on coarse-grained models. His research group investigated whether small peptide mixtures could effectively capture fundamental properties such as fluidity and aggregation propensity, employing a "sticker-spacer" framework where glycine/serine spacers were combined with tyrosine stickers. These ternary mixture simulations successfully reproduced phase separation within elongated boxes while maintaining smoother density profiles that indicated the presence of condensate fluidity, offering an alternative perspective to pure sticker systems often used in broader studies.
A central theme of the presentation was resolving discrepancies between chemical intuition and experimental observations regarding the "stickiness" of aromatic residues, specifically comparing tyrosine and phenylalanine. Although hydrophobicity scales typically suggest that phenylalanine is stickier than tyrosine, experiments demonstrated that tyrosine induces phase separation more efficiently with a lower saturation concentration. To explain this counterintuitive result, Sancho's team used alchemical transformation methods within thermodynamic cycles and quantum chemical calculations to measure transfer free energy differences across varying dielectric constants. They identified a crucial crossover in interaction strengths where tyrosine is favored in the intermediate dielectric environments typical of biomolecular condensates, whereas phenylalanine becomes favorable only in low-dielectric folded protein cores or non-hydrogen-bonding solvents.
The study also examined positively charged residues like arginine and lysine, confirming through simulations that arginine has a higher propensity to transfer into condensates due to the high desolvation penalty associated with lysine, a hierarchy consistent across different solvent environments. The presentation highlighted significant force field dependence in these results; while older models like ff99SBILDN performed well, others such as ff14SB-disp failed to reproduce phase separation in ternary mixtures, underscoring the necessity of careful validation against experimental data when selecting simulation parameters and water models. Furthermore, researchers tested whether interaction propensities were driven solely by tyrosine-tyrosine interactions or influenced by other factors; swapping tyrosine for phenylalanine showed that while pi-stacking might be expected with the latter, overall interaction patterns remained largely unaffected in validated pentapeptide test beds using GGXGG sequences.
In conclusion, the research established that aromatic and charged patterning interactions are complementary rather than competitive when positively charged residues exist within aromatic-rich environments, distinguishing these mechanisms from coacervates formed purely by electrostatic attraction between opposite charges without aromatics. The findings suggest that despite computational limits currently favoring smaller models over larger systems found in literature, the consistency across different force fields indicates robustness for using these proxy studies to investigate full-length intrinsically disordered regions. This work not only clarifies the fundamental drivers of condensate formation but also sets the stage for future investigations into conformational ensembles and multi-scale modeling approaches that integrate drug design and artificial intelligence in biological research.
Read the full video transcript
Welcome everybody
to the BioExcel webinar number 97.
Today, David Sancho's the Sancho, sorry,
will
speak about driving force in
biomolecular condensate from atomistic
simulation of model peptides.
David is from the University of the
Basque Country,
>> [snorts]
>> but is also affiliated to Donostia
International Physics Center.
I'm Alessandra Villa from the BioExcel
Center of Excellence for Computational
Biomolecular Research, and I host this
webinar together with Otto and Richard.
I want to inform everybody that the
webinar is recorded.
During the webinar, you can ask question
using the Q&A
function that you find at the bottom of
the Zoom application.
Depending on the operating system that
you have, you can see one of those three
symbols.
You can just click it and type your
question.
So, I know we know that you have a
question, and we will read the question
at the end of the webinar.
After the webinar, maybe you still have
some question,
so you're most welcome to join us in the
Ask BioExcel forum.
Just you can use the barcode, or you
just click on the link,
and you go to the
post on the webinar of David, and you
can ask your question, and David is
available to answer the question for the
next week.
So, now something about the presenter of
today. So David is professor at the
University of the Basque Country
and on the Donostia International
Physics Center.
He received a trainer in 2023
uh following to a Ramon y Cajal
fellowship.
Actually his journey in this field start
with a PhD on genetic algorithm and
coarse-grained model. After that he move
to the Spanish Research Council as a
postdoc and to the University of
Chemistry of Cambridge where he start to
work on molecular simulation.
Currently is the leader of the
biomolecular theoretical chemistry group
where he dive into the complex world of
intrinsic disordered protein using mix
of computational chemistry and
biomolecular simulation.
And today he will speak us about the
driving force.
So I stop sharing and I give the
opportunity
>> [snorts]
>> to David to share his screen. Please.
>> Uh thank you very much
Alessandra for the
kind introduction and for also for the
opportunity to present my work in this
venue.
I've
participated in the BioExcel community
for the last couple of years and it's
been a a very good experience to be part
of it as an ambassador. Um
the group I'm part of is a large
community of researchers based on this
beautiful city of San Sebastian
where we
undertake computational
chemistry work in many different areas
including polymer development of
methods, sustainable chemistry and then
there's a small group of people
who are focused on biological systems.
These are the current members of the
group. There's four
permanent members,
which are here featured on the top,
including Xavier Lopez,
myself, and Elena Formoso and Jon
Uranga, and a number of postdocs and
students that are part of
the laboratory. And together we focus on
the more biological aspects of
computational chemistry. And to in
today's talk, as Alessandra said, I'll
be speaking about our recent work on
biomolecular condensates that we have
been approaching by studying model
systems, model peptides.
I will start by discussing how this is a
general approach that actually we've
been
trying to participate on
for like well multiple
different problems. I will then focus on
phase separation and provide this
audience, which is more on the
biomolecular simulation methods than
maybe on this type of biological
problem,
well some fundamental facts about how
experiments and simulations intermingle
to try to understand
biomolecular phase separation better.
And then I will show how we've been
using peptide
models to try to understand
these condensates for a special type of
protein, the low complexity domains,
which are a type of intrinsically
disordered protein. And finally, I will
explain how we've been using
actually some bioexcel tools for
molecular simulation to try to
understand uh, what I will define as
sticker strengths, but uh, more on that
uh, later. Uh, I will start by telling
you about this general idea of using uh,
peptide models to study complicated
problems like phase separation. Uh, I
guess I'm a protein folder at heart and
uh, when trying to understand this
problem, the problem of uh, protein
folding, uh, small systems have been
instrumental for the community to
understand well, all the uh, mechanisms
of these uh, complex uh, reaction
involving the rearrangement of many
degrees of freedom. Uh, actually in my
own work, I've been looking at these as
small bits, as small elements of
secondary structure like hairpins or
alpha helices as model systems for the
more complicated uh, process and and a
particular example comes from when the
protein folding community considered
that trying to understand what uh,
something called internal friction that
characterizes folding reactions uh,
actually was uh, because its origin was
rather
uh, elusive and by looking at something
as simple a toy model like the alanine
dipeptide, uh, what experimentalists
called internal friction which as it was
an insensitivity to the viscosity of the
of the folding times happened to be uh,
existing even in these uh, simplest
system that undergoes a conformational
transition between uh, different um,
states.
So, this uh, reductionist approach of
ours is something that we apply in
different
uh, well, aspects, in different uh,
types of problems and some recent work
is for example when we were uh, trying
to understand neurofilament peptides
that seem to uh, aggregate in the
presence of metals. Specifically, there
are some some uh, fragments that happen
to be particularly important for this.
And by looking at the sequences, we we
identified that there were a number of
repeats that we simulated independently
to see how these small
peptides were already able to capture
some special properties of the of the
neurofilaments.
Another recent example comes from our
work in A beta, which is involved in
Alzheimer's and which forms these famous
fibros.
We try to understand the system by
reducing it specifically in the case of
of metal binding again to smaller
fragments that we could characterize
in detail
using ensemble refinement methods. And
another example, a third snapshot from
our recent work, would come from these
work we've recently undertaken in
collaboration with Joseph Rogers,
where he was suggesting that there could
be some
ways of trying to hijack this
interaction. He passed on a number of
sequences from peptides having macro
cycles that were particularly efficient
in these task. And we ended up where we
typically
find ourselves in in the Ramachandran
map on how small
bits of structure formation were
determining the the global properties of
these of these systems.
But now I'll shift gears and tell you a
little bit about membraneless
membraneless organelles,
which maybe this community is not so
acquainted that they have recently
attracted a lot of attention as a new
paradigm of cellular organization.
I think in this cartoon you will see
that in addition to the usual membrane
bound organelles,
in this representation of the cell we
see lots of additional cytoplasmic
actually nucleolar granules that happen
to be formed by biomolecules that
without the need of a membrane
condense and form
these organelles and well, lots of roles
have been attributed to these
type of of granules and and for this
reason we have also been paying
attention to what's going on with them.
The community has used the language of
phase transitions and then
been able to produce these type of phase
diagrams of temperature versus
concentration where at a given range of
concentrations you're able to find these
type of
two phase
regime and in these
regime we find the formation of
biomolecular
condensates.
An additional reason why these
condensates are relevant is because they
have
been recognized to form to be precursors
of
solid aggregates that can be
pathological and this is something that
we are finding in many many different
systems
well, involved in in neurodegenerative
disease. So these systems these
biomolecular condensates are not only
interesting for their physics but also
for their
biomedical relevance. And if we think on
the types of proteins that undergo these
type of transitions, we could find many
different types of proteins, folded
proteins,
linear multivalent proteins but I will
focus my discussion primarily on
something that we spend a lot of time
with, intrinsically disordered proteins
that seem to be having a special
residues, adhesive residues that
determine the phase separation when they
are at high concentrations.
And when the community has to try to
understand these systems
computationally, well, they turn out to
be quite challenging because these as
you as the
image suggests, these
proteins occupy large conformational
ensembles. They are occupy large volumes
and of course running simulations on
these
IDPs, intrinsically disordered proteins
is quite
complicated.
So typically,
the community has resorted instead
going for coarse-grained simulations
that have been probably the most useful
tool for understanding biomolecular
phase separation from a computational
standpoint. This is work by Robert Best,
Jitain Li,
Greg Dignon and Wenwei Zeng who were
running slab simulations of a very
simplified model where that allows for
running simulations of mixtures of many
different proteins that will form
like well condensate at the center of
these pieces lab
and then eventually will be able to to
escape from the from the condensate
every now and then. If the same
simulation is run at a different
temperature, at higher temperature, you
will see that the
that the condensate dissolves and by
doing this many different temperatures,
you are able to compute
the
density profiles
where you see that at the low
temperature, you have a condensate at
the center and the density of the
condensate gradually decreases as you
increase the temperature
until you eventually reach the single
phase
regime with these flat density profile.
From the concentration of the dilute
phase and the dense phase, then you're
able to compute uh these uh points, the
saturation concentrations at different
temperatures and the dense phase
concentrations, and from that you're
able to obtain the phase diagram of the
system with
characteristic parameters like the
critical concentration and um
temperature.
This uh
avenue of uh molecular simulation based
on coarse-grained models has been
extremely successful in the last decade
in the study of uh biomolecular
condensates. Typically, uh
one starts from from one has to start by
thinking about statistical potentials
that are usually the base of some of the
of the um
energy functions used in these um
in these type of simulations. Uh I guess
uh methods developed by Gerhard Hummer
and John King were also instrumental for
being able to construct these types of
models, and in 2018, Jeetain Mittal and
his coworkers put together the HPS
model, that is a real really
foundational model for a lot of work
that has come afterwards. There's a
large zoo of uh models of coarse-grained
models that one can now use with
increasing accuracy, being able to
capture all sorts of details based on
the protein sequences and also
post-translational uh modifications.
So, this has This avenue has been
extremely successful,
but of course it misses some details
that we could aim to capture using
atomistic molecular dynamics. And uh one
thing that one can do is try to do these
uh atomistic simulations of uh similar
systems. Uh this is something that has
been attempted in the past uh by a
number of people. Uh this is again work
by Jeetain Mittal and Robert Best and
their coworkers, where
what what what what what one normally
needs to do is start from the
coarse-grained simulation, then rebuild
the full atomistic resolution, then sort
of re-equilibrate the the system
solvated, and
again just be able to run the atomistic
MD. But, in these types of
circumstances, one misses a lot of
things.
It's typically very difficult to run
these simulations, and even in in
microsecond time scales, you wouldn't be
able to ever sample the formation of the
condensate and many other
interesting
details of the of the system. That
hasn't precluded the community to engage
into these massive
scale simulations, where you can see
that actually,
well, one is able to to model and
simulate the condensate for many chains
of different types of proteins. But, as
I say, this comes with a a number of
complications
in particularly in the setup and of
course in the running of the
of the calculations. So, I guess that
our reductionist
idea is to try to see how small a system
can capture some of these
properties as well. And specifically,
this is something that has been done by
others in other contexts in the past.
So, it's not that I'm claiming their
idea is original
to us.
But, I I'll tell you what we've been
working on specifically.
We started thinking about a special type
of protein,
some intrinsically disordered
regions of proteins called LCRs due to
their low sequence complexity. And
what you see here in this slide is well
a number of usual suspects in the case
of phase separation. These have been
characterized experimentally by many
research teams. And when one explores
these sequences, what you immediately
see is that their composition is very
very surprising. At least if you are
used to the sequences of proteins that
are able to fold. You will find long
stretches of glycine amino acids like
here and here. You will find lots of
polar residues like serine. And also you
will see that these sequences seem to be
interspersed by automatic and positively
charged amino acid residues like
phenylalanine, tyrosine, or arginine
that seem to be characteristic of these
proteins and that seem to be involved in
these
phase separation.
So
in combined experimental and theoretical
work,
people like Tony Mitach and Rohit Pappu
have rescued these polymer model, the
sticker and spacer model that tries to
capture the peculiarities of these of
these sequences.
Basically, the sequence would be
described as a combination of spacers
and stickers. The spacers giving
flexibility and fluidity to the
condensate, the stickers providing the
addition properties and two design
parameters for the sequences being both
the multivalency and the patterning. And
as I say, this model has been extremely
successful both for exploring the single
chain properties of these disordered
proteins for these disordered regions of
proteins by showing how having more or
less aromatic residues, more or less of
these stickers, you're going to be
altering the dimensions and the polymer
scaling principles or the polymer scale
scaling powers of the of the
proteins at hand. But in addition, these
number of stickers, number of aromatics,
for example, in this case are able to
predictive of the phase diagrams for
these types of systems. And
additionally, they seem to be
determining the aggregation propensity
of the system. So when you class cluster
too many of these aromatic residues in a
given sequence, instead of forming
well-formed spherical droplets, you fall
into a more aggregation-like
regime
as in these case.
So, inspired by all of this work, we
thought that maybe something very very
simple
like combining only
well, a subset of amino acid residues
with composition inspired by what we
knew from these low complexity regions.
Basically, by getting just a few
ingredients, just a few amino acids, and
putting them into a simulation box, we
would be able to capture some of the
properties of these of these
condensates. And as it turned out, we
were able to to to obtain this type of
phase separation in in elongated boxes
like the ones that they use in the
coarse-grained models.
Just some technical details for for the
well, people who are interested. We're
not doing something
nothing like particularly exotic here.
We're generating our peptides with amber
tools. We're packing our boxes
with packmol and we use bioexcel's
favorite simulation
tool gromacs to run our simulations with
a workflow that doesn't differ much from
the standard solvation, addition of
ions, minimization, equilibration, and
um production
uh stages.
Uh of course, we need to pay attention
to details like which force fields we're
using. We have used relatively old force
fields that had been optimized by other
people like 99 SBILDN or FF03* from the
Amber family. We've also been trying
with this uh more recent force field
from the
uh Robustelli and his co-workers.
And uh we've used uh different types of
uh water models, but of course, there's
a lot to explore in terms of uh force
fields.
Uh I probably should mention that
specific force fields that we are
currently using have been developed uh
specifically uh with a focus on uh
biomolecular uh condensation, so
probably that's the one that one should
be using in these uh type of um
the study.
But back to what we did, uh we started
by just running simulations of uh
spacers. So, these would be glycine and
serine. Uh just adding glycine into our
simulation boxes, we would find that uh
there was no phase separation, and and
when we
uh used these uh uh
obtained these density profiles, we
would see that we were a flat density
profile uh
for this uh spacer-only uh simulation
box.
Uh the same happened uh when we were
running simulations only for serine,
uh but we obtained something very
different when we run simulations for
the uh speaker residue. These are
results for uh tyrosine, and in this
case, we would obtain a density profile
with a hump in the middle. So, uh of
course, tyrosine, which is much
stickier, would be aggregating and you
actually can see that it's jagged in the
center. That's because even if we're
averaging many frames for obtaining
these density profiles,
I guess the the dynamics are extremely
extremely
slow.
So, we're capturing different behavior
with these types of simulations, which
is essentially what we intended to do.
We then went into mixtures and we first
did a a mixture of
spacers. And again, we obtained the flat
density profiles, but when we put the
three components in the simulation box,
we obtain an interesting result, which
is that again, we obtained phase
separation, but now with much smoother
density profiles, suggesting that this
ternary mixture actually was able to
preserve some of the fluidity that has
had been the case of of tyrosine. So, we
produced simulations at many different
concentrations and we found that the
results that we had obtained seemed to
hold together pretty well. This is for
the single component simulation boxes
and also for the binary and ternary
mixtures. So, the results seem to be
quite robust to a protein compositions.
We looked at some parameters that are
typical in these in these type of study,
like how much the accessible surface
area or the
the size of the of the largest cluster
would change. You can probably see here
that as a function of time, we see a
very different behavior from for the
sticker than we do for the spacers. And
then the ternary mixture is somewhere in
between. That's true for the accessible
surface area and for the size of the
clusters. You can compute some metrics
called aggregation propensity and
clustering degree. Here we're
aggregating simulations at different
concentrations, and you will always find
that the ternary mixture is um somewhere
in between.
Uh, so we thought that this was maybe
capturing something that's true also for
the ID uh for the IDPs.
We obtained a phase diagram by running
temperature replica exchange uh
simulations, and again, here we're
showing uh density profiles, with these
being the protein density profile and
that of uh water, which of course is
like a a sort of mirror image. At high
temperature, we dissolve the condensate,
and at low temperature, we get the
density profile with a very well-defined
uh peak. And from this, we're able to
compute a phase diagram.
But one further question would be
whether these uh peptide mixtures
actually resemble the larger the longer
protein condensates at all, or whether
well uh uh
similarities are just a coincidence. So,
we looked at the contact patterns uh for
these uh systems. We saw how many of
these tyrosine-tyrosine contacts there
would be, or how many tyrosine-serine
and all the possible combinations in our
simulation boxes, which were the
residues that were contributing more to
the stability of the condensate. And by
comparing them with statistics
uh from the large simulations that I was
showing uh before, we saw that the
contact patterns actually were very much
similar in our uh very simple peptide
model systems and in the full-length
protein uh simulations. So, uh well, we
thought that the we had something uh
fundamental uh right in our uh model
systems, and that since we had a a a toy
model that could allow for us to probe
uh many different
questions. We would try to find a
question that we could address using our
model peptide condensates.
And one question that the community has
been thinking about is what is the
relative sticker strength of amino acid
residues because
there are different types of stickers.
I've mentioned phenylalanine and
tyrosine. They are stickers. There's
also the positively charged ones like
lysine or arginine, but not all of them,
regardless of sharing the positively
charged or aromatic character, are
actually equally equally sticky. And
specifically we we first focused on
phenylalanine and tyrosine
because there were some very elegant
experiments again performed by Tony
Mittag and and Rohit Pappu and their
co-workers were by looking at the
sequence variants of a given protein
that had both phenylalanine and tyrosine
and by replacing phenylalanines with
tyrosines and tyrosines by
phenylalanines, they would get these
different variants. They were able to
measure a phase diagrams to obtain phase
diagrams from experiment for all of the
systems and what they found was that the
system with more tyrosines would have a
lower saturation concentration than the
system with the wild-type sequence or
even
more to an extreme, the system with many
phenylalanines. And the trend in terms
of saturation concentrations is very
clear. You need
less of the tyrosine than the
phenylalanine to induce phase
separation
indicating
quite
clearly that indeed s- that uh, indeed
uh, a a tyrosine is a stickier residue,
the one that triggers
phase separation more efficiently. And
while that uh, a phenylalanine is also
able to to trigger phase separation, you
need much more of it to uh,
these arrival into the two-phase regime
that is in these uh,
region of the of the
And
when I first looked at these results, I
was uh, surprised because this did not
quite align with what I recalled from my
PhD work about a phenylalanine and uh,
tyrosine. So, we went back to the
literature to look for some old papers
and see what our chemical intuition
would be telling us about phenylalanine
and tyrosine and their propensity to be
uh, phase separating, about their
stickiness. And the first thing we
looked at were were solvation free
energies. And in the case of uh, well,
these experimental data set for uh, side
chain analogs, you can see that the
tyrosine solvation free energy
uh, is much more favorable than that of
phenylalanine.
So, in this case, uh, I think it's quite
clear that you would expect a
phenylalanine to be the stickier residue
rather than the um,
tyrosine. Or at least, that's what we
thought.
Another piece of evidence comes from uh,
very elegant analysis carried out by
Julio de Se and Christian
Lindorff-Larsen and their co-workers,
where they were looking at many
hydrophobicity scales, uh, which were
summarized in these stickiness
parameter. And they looked at the
distribution across the many
hydrophobicity scales of this stickiness
parameter for all the amino acids. And
what they found, if we focus on the two
residues uh,
that we are studying is that again
phenylalanine would be the more
hydrophobic
residue or the more sticky residue while
tyrosine would appear at much lower
values on average. So again pointing to
the fact to the expectation that
phenylalanine would be the stickier
residue contrary to the experiments.
Then we went into
statistical contact matrices
that have been
have been from from
experimental data sets of protein
structures. You can go into the PDB
retrieve lots of structures and look at
the frequency of contacts for
different types of pairs which would be
captured by these NIJ and using a
suitable references state you can
convert this frequency of contact
formation for different pairs into some
pseudo energies.
If you look at the most famous of these
statistical contact matrices for
phenylalanine and tyrosine you would
again find that from this type of
arguments you'd expect phenylalanine to
be the stickier residue rather than what
was finding the what was found in the
experiment which is telling us that
tyrosine is is much stickier than
phenylalanine.
There's also some information suggesting
that tyrosine tyrosine interactions are
stronger than those of phenylalanine
phenylalanine. You can look at
potentials of mean force like these
calculated by Colipardos Rosana
Colipardos
team and you'll see that the tyrosine
tyrosine PMF is deeper than that of
phenylalanine phenylalanine so that's
evidence pointing in the same direction
as the experimental work on condensates.
And also
you can go into these even more
enigmatic
data sets on solubilities where you will
find that in fact tyrosine appears over
an order of magnitude below in
solubility than any other residue or for
that matter
phenylalanine, again pointing to a
greater stickiness of tyrosine. So,
depending on what you look at, you'll
find that either phenylalanine or
tyrosine is a stickier residue based on
naive expectations.
So,
we wanted to go a little bit beyond
these naive ideas or our chemical
intuition from other published data
sets, and we thought that maybe what
could capture these greater propensity
of different amino acids to phase
separate could be
modeled using free energy calculations
that are used extensively by our
community, uh e- particularly when
you're trying to study
well, you're doing drug discovery
studies and you want to look at whether
one molecule or another molecule will be
binding more effectively
to a partner and by virtue of a
thermodynamic cycle, you can do an
alchemical transformation
both in both the states in both the free
and bound state and indirectly get the
delta delta G binding and and you do
this by invoking non-physical
intermediates that are
easy to simulate. So, we thought we
would do this for condensates and
specifically what we wanted to measure
in our calculations was the different
propensity of transfer of a
phenylalanine peptide into a given
biomolecular condensate or tyrosine
being transferred into the biomolecular
condensate. And we thought that maybe if
we were getting that the transfer of
tyrosine was more favorable than that of
phenylalanine, then we would be
capturing this greater greater
propensity of um
tyrosine to phase separate that had been
found in the experiments.
Uh again, we used one of the BioExcel
tools to run these calculations.
Specifically, we used uh the the very
elegant uh PMX non-equilibrium
alchemical uh switching method by uh
Bert de Groot and Vidutas Gapsys and
Mateus Delli and and their co-workers,
uh where one typically requires uh two
different uh simulations for uh both
lambda states that one is considering.
In our case, one will be phenylalanine
and the other uh will be tyrosine. And
after running these long equilibrium
simulations, you will do uh many
non-equilibrium switching runs uh both
in the forward and reverse direction
from lambda zero to lambda one and from
lambda one to lambda zero. And by
computing the distributions of work for
the forward and reverse tran-
non-equilibrium transitions and using uh
non-equilibrium theorems, one is able to
determine the uh delta G of interest.
How does this look in the case of uh
condensates? Well, in the case of
condensates, we have to first set up a
simulation box. What I'm showing here is
the density across the simulation box,
and you see that there's a dense region
and a dilute region. So, here we have
the phase separated system, and this red
line in the center is the peptide that
we are going to alchemically transform.
This we did at lambda zero for
phenylalanine peptide and lambda one for
the uh tyrosine peptide. And then we run
the switching
runs. And this type of calculation can
be done in the condensate, like in the
case of this
plot here or this other plot here, but
it can also be performed in water. So,
from those calculations, one will be
able to obtain all the delta Gs that can
then be combined to obtain the
difference in transfer free energy
between phenylalanine and and tyrosine
for these model systems.
And this is how our simulations look
like.
We have the condensate at the center. We
have lots of water molecules in the
dilute phase that sometimes you'll see
that the peptides exchange with the
dilute phase. The
faint red
in the center, that's our alchemically
transformed
amino acid residue which is at the
middle of a slightly longer peptide than
the ones that form the condensate.
So, we did these very many times for
multiple condensates. And what we
obtained was that when the condensate
was formed by glycine, serine, and
tyrosine, the delta delta G transfer was
negative, telling us that the transfer
of tyrosine, the tyrosine bearing
residue, was more favorable than the
phenylalanine bearing residue. So, in
good agreement with the experimental
work. We replaced tyrosine in the
condensate for phenylalanine to see
whether the tyrosine in the condensate
was responsible for this result, but we
found that for a condensate formed by
glycine, serine, and phenylalanine, the
result was very much
the same.
So, we thought that probably it was the
interactions that this tyrosine was
making, which would be different from
the ones that phenylalanine would be
making. And we looked at the interaction
patterns in this central region of the
protein, so of the condensate we we run
additional simulations of dense phases
alone. Uh we looked at uh statistics of
contact formation for tyrosine and
phenylalanine. We found them to be
essentially indistinguishable regardless
of which was the residue that was part
of the condensate. We looked at the
interaction patterns and regardless of
whether we were looking at pi stacking
or sp2 pi interactions as a function of
these angles and distances, we would
find that essentially there there was
not that much of a difference in terms
of the interaction patterns that were
determining these um result.
Uh of course, there's a main difference
uh that between phenylalanine and
tyrosine and is that tyrosine can form
hydrogen bonds,
but also one has to notice that hydrogen
bonds can also be formed outside, so it
is ambiguous whether this is making a
strong contribution in terms of the
energetics for giving a preference to
tyrosine relative to phenylalanine. And
we thought that the condensates were in
some way acting like uh a solvent that
was contributing a medium
uh and in this medium, uh which is
highly hydrated and is capable of of
forming hydrogen bonds, uh well, this
was making being preferable for the for
the tyrosine than it was for the
phenylalanine and this is what's
determining the greater uh propensity to
to phase separate. But if condensates
promote this trend, maybe different
solvents would promote a different
trend. So, we run these alchemical
transformations and computed these delta
delta D transfer into many other media
um like acetone and methanol or ethanol,
all of which can form uh hydrogen bonds,
but have a lower dielectric uh constant
and we found that these
negative
well this preference for tyrosine
actually was reversed as we were
increasing the size of the aliphatic
side chain of the alcohols.
So there seems to be a crossover between
a tyrosine preferring and a
phenylalanine preferring region and this
is accentuated even more
when you go into really a polar solvent
where the delta delta D transfer favors
clearly the the phenylalanine relative
to relative to tyrosine. And by looking
at the dependence of this delta delta D
transfer with the
with the dielectric constant we found
that
the dielectric constant seemed to track
this crossover between the phenylalanine
favoring region and the tyrosine
preferring region.
So after
well obtaining these results my
colleague Xavier Lopez
devised this alternative way at looking
at the same problem and proposed this a
scheme for running quantum chemical
calculations for a contact formation
thinking that it would involve the
well two amino acid residues I and J and
that it could be split in two different
steps. The first step would involve the
transfer into a medium maybe the
condensate or maybe a different solvent
and
then there would be the contact
formation step and luckily with quantum
chemical calculations we have methods
that enable for us to calculate both the
steps very accurately. So one of them
would correspond to the transfer step
that can be calculated with
SMD very accurately. and additionally
one could also try to obtain the
interaction energies and combining all
of this then what would be able to get
the contact formation propensity into
the medium that we're trying to to
study.
Now we did this types of calculations
for many different arrangements of
phenylalanine and phenylalanine and
tyrosine and tyrosine including also
their interactions with backbones that
are going to be present in the
biomolecular
condensates and can be partners of these
amino acid
residues and um
summarizing many many calculations
we represented the difference in these
contact between any combination of the
of the tyrosines of phenylalanines with
respect to our reference which would be
the most stable of the of the
phenylalanine
pairs and what we found is that at
intermediate die electrics we see again
a crossover into a regime that favors
tyrosine relative to phenylalanine. Also
this is dependent on the type of solvent
in which we are doing the transfer which
in the is indicated by the different
symbols and and depends on the
configuration
at hand specifically when the solvents
are
well don't have hydrogen bonding groups
you always get that the that the
well phenylalanine is is more
likely
in the case of alcohols or a special
type of water where we tuned the
dielectric constant you get into this
region that makes tyrosine
preferable.
So we seem to have identified a
crossover between these two regimes and
we think that these regime of low
dielectrics is actually what corresponds
to the the course of folded proteins
which are characterized by low
dielectrics while this region of
intermediate dielectrics is actually
what corresponds to the tyrosine
favoring region of biomolecular
condensates.
We've calculated these dielectric
constants for the condensates. They seem
to be in the range of 30 and 40. So, in
this case we may be in the in the
tyrosine favoring regime.
So, after settling this we thought we
could continue with our work and go into
the other
sticker strengths in those of arginine
and lysine.
Both
positively charged amino acids that can
form interactions with aromatic residues
via
cation pi
pairs and again other people have looked
into this before but we did this using
our approach using our thermodynamic
cycle, the alchemical switching
transformations from PMFs for the
transfer of either
arginine or lysine, lysine here and
arginine here into the biomolecular
condensate. And the results in this case
for this type for this pair of residues
is
different
in all cases we obtain that the transfer
is favorable for arginine again in in
good agreement
with experiments. There's no clear
dependence with the
dielectric constant and what was very
interesting that came out from these
uh, quantum chemical calculations that
we performed also with support from a
very talented undergraduate students,
uh, Lydia Armentia from our lab, is that
specifically the, uh, well, lysine seems
to be in in the very, uh, well, low
propensity to desolvate, uh, for lysine,
uh, which is, uh, well, much more
difficult to transfer into the
condensate that arginine uh, in the case
of arginine, the positive charge is
delocalized and that favors, uh, the
transfer into into the condensate. So, I
will conclude here with these, uh, last
results. Uh,
I've tried to tell you about how these,
uh, LCR-like amino acid mixtures that
we've been putting together using these,
uh, simple models
seem to, uh, capture some of the
interesting properties of intrinsically
disordered proteins that are able to
phase separate. Then I've told you about
our work of on tyrosine and
phenylalanine using, uh, alchemical
transformations and DFT calculations
where we identified a crossover in
interaction strengths that we believe
explains the different propensity to
phase separate, but also reconcile many
observations relative to the, uh,
relative, uh, strengths of interactions
inside the cores of folded proteins.
And, uh, finally, I concluded with uh,
just a few comments on our more recent
work on arginine and lysine where the
hierarchy is, uh, always, uh,
the same regardless of the environment
due to the, uh, large penalty for
desolvation of lysine.
Uh, this work has all been published in
the references I'm showing here, so
maybe there are some more, uh, more
details uh, there, uh, if you're
interested. and uh,
with this I finalize.
I would like to thank you
for your attention and also again
bioexcel for hosting the seminar.
>> Thank you very much, David.
And we have a couple of question
in the chapter, so I will read it.
So,
the first one is how did you calculate
the aggregation propensity?
This was at roughly the middle of your
presentation.
>> Sure, sure.
Yeah, so thank you for the question.
Yeah, the aggregation propensity
and
I cannot exactly remember what the other
parameter was. Yeah, let me just go
back. These are obtained from the two
measurements that I showed first. So,
you can calculate the the change in the
accessible surface area from the
perfectly dilute
regime and also you can calculate
the maximum the maximum
cluster size. So, from the cluster size
you get the cluster what's called the
clustering degree.
And
the aggregation propensity comes from
how much you shrink
across your simulation from the
perfectly dilute state. So, in some
cases you will see that there's a large
decrease in the solvent accessible
surface area. That is very much the case
of the tyrosine which excludes all the
water and then you get a large reduction
in the accessible surface area. That
results in a high aggregation
propensity. For other systems, you get
very much
you stay at very much the same
accessible surface area because they
remain perfectly dilute and in that case
you have a very low aggregation
propensity. These are not parameters
that I have defined Uh, myself. And if
you go to the Biophysical Journal
article, you'll see the origin for these
definitions that were established by
other people, not by me.
>> Thank you.
And then there is another question
that it was was say thank you, really
nice talk. Uh
it was wondering that if your results
were consistent across force field.
And
or there is one that perform better than
another referred to force field. And
then also go on how the water model can
affect the real behavior of the
condensate.
If so, in your opinion.
>> Yeah, so I guess I should start by
saying that we haven't done a systematic
study of different force fields
in our work. I think this is actually
something that may be extremely useful
because I I guess that one would be able
to get these calculations
quite rapidly and and and across force
field. So maybe it's a it's a good place
for doing a force field validation
exercise
that we really haven't done. We have
done some tests though and
I guess I I started by by mentioning
that we've
used the
our work primarily relatively old force
fields. So much of our work has been
done with FF99 SB12DN.
Uh we tried the same systems
with FF03 star with backbone
corrections. It yielded very much the
same results.
But then we obtain a very different
result in the case of the 99 SB disp
force field which also comes with a
different water model.
Uh
in this case, we wouldn't see phase
separation for the ternary system. Uh
this is something that was not too
concerning for us because I guess some
deficiencies in binding properties have
been
identified for these 1999 1999 SB disp
force field. So, in that respect, maybe
this this is a result that we could
expect.
Um we're doing now work with
uh force fields optimized by Jain et al.
precisely for the study of condensates.
I would say there is going to be
definitely dependence on the models that
you are using and maybe the force fields
that we have been using and with the
water model, the tip 3p water model,
they are more collapse-inducing.
Uh
our recent work and the work of many
others on the small model systems
seems to indicate that with a degree of
variation,
uh well, there is consistency across the
board that the type of small peptide
models that we're using, well, seem to
be
forming these types of
condensates, but I would agree that this
is probably something that requires
careful validation against experiment.
>> Thank you. So, we have another two
question. So, the first one is where the
simulation performed for
GSF or G GFY amino acid mixtures?
And then there is a follow-up, did you
use any peptide? If so, what length?
And would there be any limit
to the length of the peptide that can be
simulated in this way with this
approach.
I guess.
>> Yeah, so I guess there's a number of
different things there on the GSF GSY
part.
We did both.
So [snorts]
in the in the first
paper in biophysical journal,
we only reported like the ternary
mixture with the tyrosine
well as an as an ingredient of the
ternary mixtures. In the more recent
work from last year,
we did both the GSF and the GSY and the
reason was not so much that we were
interested in the different mixture as
such but on the possibility that the
propensity we were finding for tyrosine
was due to it having other tyrosines to
interact with. So we wanted to swap the
component of the mixture to see whether
interaction patterns would be
drastically affected with maybe much
more pi stacking in the case of
phenylalanine
but that was not the case. So we did
both GSF and GSY.
Then I think there was
another part of the question which
related
related to the size of the peptide.
>> Well, I guess
>> And then 80 amp used one was asking
which length.
>> Yes, so so I guess that we are using the
dipeptide or the terminally terminally
blocked
residue monomers
for the condensate but the alchemical
transformation is performed on a
pentapeptide
with an residue that changes the residue
that is alchemically transformed
is the one that can either be
phenylalanine or can be
tyrosine. This is a peptide that has
been this pentapeptide has been studied
using NMR and and many different uh
methods in the past. So, I guess it's a
good test bed. It simply has a sequence
GGXGG.
So, it's just glycines and a host
residue at the center or the guest
rather residue at the
at the center.
And in all cases, the peptides are
terminally blocked. So, both for the
amino acid mixtures and for the peptide
we are alchemically transforming
transformed they are not zwitterionic.
They are uh
neutral terminally capped uh residues.
And I think there was one one more
thing, no, Alessandro?
>> Yeah, so but you mute one. You try you
alchemical transformation is involving
one residue.
>> Yes. Yes. Just one.
>> Okay. Yeah. In in a pentapeptide?
>> In a pentapeptide.
>> Yeah. And then the last question was
if there are any limit in the length of
the peptide which for this type of
simulation.
>> Yeah.
So, I guess uh
well, in terms of
limitations could come in in in in
different ways.
Uh I guess that one will be on based on
computing power and probably our
computing power will end up
uh early. So, that's why we go small and
and decide to to study
like these uh uh
little
uh systems. Um when I was before in a
different question claiming some
robustness in the results regardless of
force fields is because we're not the
only ones doing this. Uh people have
been looking at these peptide
condensates for systems that I would say
are up to some eight residues. Uh
you should look at the work by David
Botoyan, Daryl Joseph, and uh many other
uh people who are looking at these types
of condensate peptide uh
models
um as a proxy for the bigger systems.
and we seem to all be uh getting
consistent results. Tetrapeptides have
been done a lot
at the beginning of
I talked I was showing simulations of
full length IDRs for proteins of over
100 residues. So, depending on your
computing power and the
difficulties in setting up the
simulations that I mentioned at the
beginning,
I guess you can simulate as far as you
can.
>> Yeah.
Thank you. I I think we we still have
two questions, but I take only one for
reason of time and I will ask you to
briefly answer because we are running
out of time. Thank you.
And
so, they thank this the attendee thanks
for the great talk. As you have nicely
shown, thickness
induced by tyrosine and phenylalanine is
one of the reason of the formation of
the condensate.
And
other is charged residue and more
precisely charged patterning.
What is in your opinion
on the relationship between the two?
Do you think that a thick patcher
compete with charged patches or on the
contrary
they are complement or complete
with one of with the other
in the disorder
in intrinsic disorder protein?
>> Yeah, so that's a a great question and
one that probably would
be better answered by some people who've
looked at the sequences in more detail.
I think that an important point is that
there's different types of sequences
that are involved in phase separation.
What we've been looking at are aromatic
driven and also cation-pi
driven. So, in our case, when we're
looking at the charged positively
charged residues, we're looking at the
interaction of these positively charged
with a condensate that has lots of
aromatics. In this case, would be
complementarity,
not competition.
But, I guess I should also mention that
there's also biomolecular coacervates
that are
induced by charged positively charged
with negatively charged amino acid
residues. And in that case, you don't
really need the aromatics. So, I guess
condensates come in in all sorts of
shapes and and forms. And I'm not sure
there's much of a competition
between both
types of possibilities. I think they're
just different.
>> Okay, thank you very much. I will ask
kindly the last question to if the
attendees can post it on the forum.
And then we I thank you a lot to David.
I just do a small announcement and then
we close this session.
So, I want to just to tell you about the
next webinar and the last for this
spring
season. And that is the 26th of May
where Kerstin Lindorff-Larsen will speak
about conformation ensemble of
intrinsically disordered region and
protein.
And last not least, we have postponed
the abstract deadline of our
weekly bi-yearly conference that will be
the
31st of May. So, there are the deadline
of abstract the conference will be in
September.
And we'll cover bio
uh sorry, multi-scale modeling,
molecular dynamics, free energy
calculation, integrative modeling, drug
design, and AI. So, please come so we
all meet in person. It will be great.
So, don't miss the deadline.
Thank you very much again to everybody,
and we see in two weeks. Bye-bye.