Submind YouTube summaries
Thumbnail for Bioexcel webinar #97: Driving forces in biomolecular condensates from atomistic simulations

Bioexcel webinar #97: Driving forces in biomolecular condensates from atomistic simulations

Watch on YouTube

Video summary

David Sancho from the University of the Basque Country presented a reductionist approach to understanding biomolecular condensates by utilizing atomistic simulations on model peptides rather than relying solely on coarse-grained models. His research group investigated whether small peptide mixtures could effectively capture fundamental properties such as fluidity and aggregation propensity, employing a "sticker-spacer" framework where glycine/serine spacers were combined with tyrosine stickers. These ternary mixture simulations successfully reproduced phase separation within elongated boxes while maintaining smoother density profiles that indicated the presence of condensate fluidity, offering an alternative perspective to pure sticker systems often used in broader studies. A central theme of the presentation was resolving discrepancies between chemical intuition and experimental observations regarding the "stickiness" of aromatic residues, specifically comparing tyrosine and phenylalanine. Although hydrophobicity scales typically suggest that phenylalanine is stickier than tyrosine, experiments demonstrated that tyrosine induces phase separation more efficiently with a lower saturation concentration. To explain this counterintuitive result, Sancho's team used alchemical transformation methods within thermodynamic cycles and quantum chemical calculations to measure transfer free energy differences across varying dielectric constants. They identified a crucial crossover in interaction strengths where tyrosine is favored in the intermediate dielectric environments typical of biomolecular condensates, whereas phenylalanine becomes favorable only in low-dielectric folded protein cores or non-hydrogen-bonding solvents. The study also examined positively charged residues like arginine and lysine, confirming through simulations that arginine has a higher propensity to transfer into condensates due to the high desolvation penalty associated with lysine, a hierarchy consistent across different solvent environments. The presentation highlighted significant force field dependence in these results; while older models like ff99SBILDN performed well, others such as ff14SB-disp failed to reproduce phase separation in ternary mixtures, underscoring the necessity of careful validation against experimental data when selecting simulation parameters and water models. Furthermore, researchers tested whether interaction propensities were driven solely by tyrosine-tyrosine interactions or influenced by other factors; swapping tyrosine for phenylalanine showed that while pi-stacking might be expected with the latter, overall interaction patterns remained largely unaffected in validated pentapeptide test beds using GGXGG sequences. In conclusion, the research established that aromatic and charged patterning interactions are complementary rather than competitive when positively charged residues exist within aromatic-rich environments, distinguishing these mechanisms from coacervates formed purely by electrostatic attraction between opposite charges without aromatics. The findings suggest that despite computational limits currently favoring smaller models over larger systems found in literature, the consistency across different force fields indicates robustness for using these proxy studies to investigate full-length intrinsically disordered regions. This work not only clarifies the fundamental drivers of condensate formation but also sets the stage for future investigations into conformational ensembles and multi-scale modeling approaches that integrate drug design and artificial intelligence in biological research.
Read the full video transcript
Welcome everybody to the BioExcel webinar number 97. Today, David Sancho's the Sancho, sorry, will speak about driving force in biomolecular condensate from atomistic simulation of model peptides. David is from the University of the Basque Country, >> [snorts] >> but is also affiliated to Donostia International Physics Center. I'm Alessandra Villa from the BioExcel Center of Excellence for Computational Biomolecular Research, and I host this webinar together with Otto and Richard. I want to inform everybody that the webinar is recorded. During the webinar, you can ask question using the Q&A function that you find at the bottom of the Zoom application. Depending on the operating system that you have, you can see one of those three symbols. You can just click it and type your question. So, I know we know that you have a question, and we will read the question at the end of the webinar. After the webinar, maybe you still have some question, so you're most welcome to join us in the Ask BioExcel forum. Just you can use the barcode, or you just click on the link, and you go to the post on the webinar of David, and you can ask your question, and David is available to answer the question for the next week. So, now something about the presenter of today. So David is professor at the University of the Basque Country and on the Donostia International Physics Center. He received a trainer in 2023 uh following to a Ramon y Cajal fellowship. Actually his journey in this field start with a PhD on genetic algorithm and coarse-grained model. After that he move to the Spanish Research Council as a postdoc and to the University of Chemistry of Cambridge where he start to work on molecular simulation. Currently is the leader of the biomolecular theoretical chemistry group where he dive into the complex world of intrinsic disordered protein using mix of computational chemistry and biomolecular simulation. And today he will speak us about the driving force. So I stop sharing and I give the opportunity >> [snorts] >> to David to share his screen. Please. >> Uh thank you very much Alessandra for the kind introduction and for also for the opportunity to present my work in this venue. I've participated in the BioExcel community for the last couple of years and it's been a a very good experience to be part of it as an ambassador. Um the group I'm part of is a large community of researchers based on this beautiful city of San Sebastian where we undertake computational chemistry work in many different areas including polymer development of methods, sustainable chemistry and then there's a small group of people who are focused on biological systems. These are the current members of the group. There's four permanent members, which are here featured on the top, including Xavier Lopez, myself, and Elena Formoso and Jon Uranga, and a number of postdocs and students that are part of the laboratory. And together we focus on the more biological aspects of computational chemistry. And to in today's talk, as Alessandra said, I'll be speaking about our recent work on biomolecular condensates that we have been approaching by studying model systems, model peptides. I will start by discussing how this is a general approach that actually we've been trying to participate on for like well multiple different problems. I will then focus on phase separation and provide this audience, which is more on the biomolecular simulation methods than maybe on this type of biological problem, well some fundamental facts about how experiments and simulations intermingle to try to understand biomolecular phase separation better. And then I will show how we've been using peptide models to try to understand these condensates for a special type of protein, the low complexity domains, which are a type of intrinsically disordered protein. And finally, I will explain how we've been using actually some bioexcel tools for molecular simulation to try to understand uh, what I will define as sticker strengths, but uh, more on that uh, later. Uh, I will start by telling you about this general idea of using uh, peptide models to study complicated problems like phase separation. Uh, I guess I'm a protein folder at heart and uh, when trying to understand this problem, the problem of uh, protein folding, uh, small systems have been instrumental for the community to understand well, all the uh, mechanisms of these uh, complex uh, reaction involving the rearrangement of many degrees of freedom. Uh, actually in my own work, I've been looking at these as small bits, as small elements of secondary structure like hairpins or alpha helices as model systems for the more complicated uh, process and and a particular example comes from when the protein folding community considered that trying to understand what uh, something called internal friction that characterizes folding reactions uh, actually was uh, because its origin was rather uh, elusive and by looking at something as simple a toy model like the alanine dipeptide, uh, what experimentalists called internal friction which as it was an insensitivity to the viscosity of the of the folding times happened to be uh, existing even in these uh, simplest system that undergoes a conformational transition between uh, different um, states. So, this uh, reductionist approach of ours is something that we apply in different uh, well, aspects, in different uh, types of problems and some recent work is for example when we were uh, trying to understand neurofilament peptides that seem to uh, aggregate in the presence of metals. Specifically, there are some some uh, fragments that happen to be particularly important for this. And by looking at the sequences, we we identified that there were a number of repeats that we simulated independently to see how these small peptides were already able to capture some special properties of the of the neurofilaments. Another recent example comes from our work in A beta, which is involved in Alzheimer's and which forms these famous fibros. We try to understand the system by reducing it specifically in the case of of metal binding again to smaller fragments that we could characterize in detail using ensemble refinement methods. And another example, a third snapshot from our recent work, would come from these work we've recently undertaken in collaboration with Joseph Rogers, where he was suggesting that there could be some ways of trying to hijack this interaction. He passed on a number of sequences from peptides having macro cycles that were particularly efficient in these task. And we ended up where we typically find ourselves in in the Ramachandran map on how small bits of structure formation were determining the the global properties of these of these systems. But now I'll shift gears and tell you a little bit about membraneless membraneless organelles, which maybe this community is not so acquainted that they have recently attracted a lot of attention as a new paradigm of cellular organization. I think in this cartoon you will see that in addition to the usual membrane bound organelles, in this representation of the cell we see lots of additional cytoplasmic actually nucleolar granules that happen to be formed by biomolecules that without the need of a membrane condense and form these organelles and well, lots of roles have been attributed to these type of of granules and and for this reason we have also been paying attention to what's going on with them. The community has used the language of phase transitions and then been able to produce these type of phase diagrams of temperature versus concentration where at a given range of concentrations you're able to find these type of two phase regime and in these regime we find the formation of biomolecular condensates. An additional reason why these condensates are relevant is because they have been recognized to form to be precursors of solid aggregates that can be pathological and this is something that we are finding in many many different systems well, involved in in neurodegenerative disease. So these systems these biomolecular condensates are not only interesting for their physics but also for their biomedical relevance. And if we think on the types of proteins that undergo these type of transitions, we could find many different types of proteins, folded proteins, linear multivalent proteins but I will focus my discussion primarily on something that we spend a lot of time with, intrinsically disordered proteins that seem to be having a special residues, adhesive residues that determine the phase separation when they are at high concentrations. And when the community has to try to understand these systems computationally, well, they turn out to be quite challenging because these as you as the image suggests, these proteins occupy large conformational ensembles. They are occupy large volumes and of course running simulations on these IDPs, intrinsically disordered proteins is quite complicated. So typically, the community has resorted instead going for coarse-grained simulations that have been probably the most useful tool for understanding biomolecular phase separation from a computational standpoint. This is work by Robert Best, Jitain Li, Greg Dignon and Wenwei Zeng who were running slab simulations of a very simplified model where that allows for running simulations of mixtures of many different proteins that will form like well condensate at the center of these pieces lab and then eventually will be able to to escape from the from the condensate every now and then. If the same simulation is run at a different temperature, at higher temperature, you will see that the that the condensate dissolves and by doing this many different temperatures, you are able to compute the density profiles where you see that at the low temperature, you have a condensate at the center and the density of the condensate gradually decreases as you increase the temperature until you eventually reach the single phase regime with these flat density profile. From the concentration of the dilute phase and the dense phase, then you're able to compute uh these uh points, the saturation concentrations at different temperatures and the dense phase concentrations, and from that you're able to obtain the phase diagram of the system with characteristic parameters like the critical concentration and um temperature. This uh avenue of uh molecular simulation based on coarse-grained models has been extremely successful in the last decade in the study of uh biomolecular condensates. Typically, uh one starts from from one has to start by thinking about statistical potentials that are usually the base of some of the of the um energy functions used in these um in these type of simulations. Uh I guess uh methods developed by Gerhard Hummer and John King were also instrumental for being able to construct these types of models, and in 2018, Jeetain Mittal and his coworkers put together the HPS model, that is a real really foundational model for a lot of work that has come afterwards. There's a large zoo of uh models of coarse-grained models that one can now use with increasing accuracy, being able to capture all sorts of details based on the protein sequences and also post-translational uh modifications. So, this has This avenue has been extremely successful, but of course it misses some details that we could aim to capture using atomistic molecular dynamics. And uh one thing that one can do is try to do these uh atomistic simulations of uh similar systems. Uh this is something that has been attempted in the past uh by a number of people. Uh this is again work by Jeetain Mittal and Robert Best and their coworkers, where what what what what what one normally needs to do is start from the coarse-grained simulation, then rebuild the full atomistic resolution, then sort of re-equilibrate the the system solvated, and again just be able to run the atomistic MD. But, in these types of circumstances, one misses a lot of things. It's typically very difficult to run these simulations, and even in in microsecond time scales, you wouldn't be able to ever sample the formation of the condensate and many other interesting details of the of the system. That hasn't precluded the community to engage into these massive scale simulations, where you can see that actually, well, one is able to to model and simulate the condensate for many chains of different types of proteins. But, as I say, this comes with a a number of complications in particularly in the setup and of course in the running of the of the calculations. So, I guess that our reductionist idea is to try to see how small a system can capture some of these properties as well. And specifically, this is something that has been done by others in other contexts in the past. So, it's not that I'm claiming their idea is original to us. But, I I'll tell you what we've been working on specifically. We started thinking about a special type of protein, some intrinsically disordered regions of proteins called LCRs due to their low sequence complexity. And what you see here in this slide is well a number of usual suspects in the case of phase separation. These have been characterized experimentally by many research teams. And when one explores these sequences, what you immediately see is that their composition is very very surprising. At least if you are used to the sequences of proteins that are able to fold. You will find long stretches of glycine amino acids like here and here. You will find lots of polar residues like serine. And also you will see that these sequences seem to be interspersed by automatic and positively charged amino acid residues like phenylalanine, tyrosine, or arginine that seem to be characteristic of these proteins and that seem to be involved in these phase separation. So in combined experimental and theoretical work, people like Tony Mitach and Rohit Pappu have rescued these polymer model, the sticker and spacer model that tries to capture the peculiarities of these of these sequences. Basically, the sequence would be described as a combination of spacers and stickers. The spacers giving flexibility and fluidity to the condensate, the stickers providing the addition properties and two design parameters for the sequences being both the multivalency and the patterning. And as I say, this model has been extremely successful both for exploring the single chain properties of these disordered proteins for these disordered regions of proteins by showing how having more or less aromatic residues, more or less of these stickers, you're going to be altering the dimensions and the polymer scaling principles or the polymer scale scaling powers of the of the proteins at hand. But in addition, these number of stickers, number of aromatics, for example, in this case are able to predictive of the phase diagrams for these types of systems. And additionally, they seem to be determining the aggregation propensity of the system. So when you class cluster too many of these aromatic residues in a given sequence, instead of forming well-formed spherical droplets, you fall into a more aggregation-like regime as in these case. So, inspired by all of this work, we thought that maybe something very very simple like combining only well, a subset of amino acid residues with composition inspired by what we knew from these low complexity regions. Basically, by getting just a few ingredients, just a few amino acids, and putting them into a simulation box, we would be able to capture some of the properties of these of these condensates. And as it turned out, we were able to to to obtain this type of phase separation in in elongated boxes like the ones that they use in the coarse-grained models. Just some technical details for for the well, people who are interested. We're not doing something nothing like particularly exotic here. We're generating our peptides with amber tools. We're packing our boxes with packmol and we use bioexcel's favorite simulation tool gromacs to run our simulations with a workflow that doesn't differ much from the standard solvation, addition of ions, minimization, equilibration, and um production uh stages. Uh of course, we need to pay attention to details like which force fields we're using. We have used relatively old force fields that had been optimized by other people like 99 SBILDN or FF03* from the Amber family. We've also been trying with this uh more recent force field from the uh Robustelli and his co-workers. And uh we've used uh different types of uh water models, but of course, there's a lot to explore in terms of uh force fields. Uh I probably should mention that specific force fields that we are currently using have been developed uh specifically uh with a focus on uh biomolecular uh condensation, so probably that's the one that one should be using in these uh type of um the study. But back to what we did, uh we started by just running simulations of uh spacers. So, these would be glycine and serine. Uh just adding glycine into our simulation boxes, we would find that uh there was no phase separation, and and when we uh used these uh uh obtained these density profiles, we would see that we were a flat density profile uh for this uh spacer-only uh simulation box. Uh the same happened uh when we were running simulations only for serine, uh but we obtained something very different when we run simulations for the uh speaker residue. These are results for uh tyrosine, and in this case, we would obtain a density profile with a hump in the middle. So, uh of course, tyrosine, which is much stickier, would be aggregating and you actually can see that it's jagged in the center. That's because even if we're averaging many frames for obtaining these density profiles, I guess the the dynamics are extremely extremely slow. So, we're capturing different behavior with these types of simulations, which is essentially what we intended to do. We then went into mixtures and we first did a a mixture of spacers. And again, we obtained the flat density profiles, but when we put the three components in the simulation box, we obtain an interesting result, which is that again, we obtained phase separation, but now with much smoother density profiles, suggesting that this ternary mixture actually was able to preserve some of the fluidity that has had been the case of of tyrosine. So, we produced simulations at many different concentrations and we found that the results that we had obtained seemed to hold together pretty well. This is for the single component simulation boxes and also for the binary and ternary mixtures. So, the results seem to be quite robust to a protein compositions. We looked at some parameters that are typical in these in these type of study, like how much the accessible surface area or the the size of the of the largest cluster would change. You can probably see here that as a function of time, we see a very different behavior from for the sticker than we do for the spacers. And then the ternary mixture is somewhere in between. That's true for the accessible surface area and for the size of the clusters. You can compute some metrics called aggregation propensity and clustering degree. Here we're aggregating simulations at different concentrations, and you will always find that the ternary mixture is um somewhere in between. Uh, so we thought that this was maybe capturing something that's true also for the ID uh for the IDPs. We obtained a phase diagram by running temperature replica exchange uh simulations, and again, here we're showing uh density profiles, with these being the protein density profile and that of uh water, which of course is like a a sort of mirror image. At high temperature, we dissolve the condensate, and at low temperature, we get the density profile with a very well-defined uh peak. And from this, we're able to compute a phase diagram. But one further question would be whether these uh peptide mixtures actually resemble the larger the longer protein condensates at all, or whether well uh uh similarities are just a coincidence. So, we looked at the contact patterns uh for these uh systems. We saw how many of these tyrosine-tyrosine contacts there would be, or how many tyrosine-serine and all the possible combinations in our simulation boxes, which were the residues that were contributing more to the stability of the condensate. And by comparing them with statistics uh from the large simulations that I was showing uh before, we saw that the contact patterns actually were very much similar in our uh very simple peptide model systems and in the full-length protein uh simulations. So, uh well, we thought that the we had something uh fundamental uh right in our uh model systems, and that since we had a a a toy model that could allow for us to probe uh many different questions. We would try to find a question that we could address using our model peptide condensates. And one question that the community has been thinking about is what is the relative sticker strength of amino acid residues because there are different types of stickers. I've mentioned phenylalanine and tyrosine. They are stickers. There's also the positively charged ones like lysine or arginine, but not all of them, regardless of sharing the positively charged or aromatic character, are actually equally equally sticky. And specifically we we first focused on phenylalanine and tyrosine because there were some very elegant experiments again performed by Tony Mittag and and Rohit Pappu and their co-workers were by looking at the sequence variants of a given protein that had both phenylalanine and tyrosine and by replacing phenylalanines with tyrosines and tyrosines by phenylalanines, they would get these different variants. They were able to measure a phase diagrams to obtain phase diagrams from experiment for all of the systems and what they found was that the system with more tyrosines would have a lower saturation concentration than the system with the wild-type sequence or even more to an extreme, the system with many phenylalanines. And the trend in terms of saturation concentrations is very clear. You need less of the tyrosine than the phenylalanine to induce phase separation indicating quite clearly that indeed s- that uh, indeed uh, a a tyrosine is a stickier residue, the one that triggers phase separation more efficiently. And while that uh, a phenylalanine is also able to to trigger phase separation, you need much more of it to uh, these arrival into the two-phase regime that is in these uh, region of the of the And when I first looked at these results, I was uh, surprised because this did not quite align with what I recalled from my PhD work about a phenylalanine and uh, tyrosine. So, we went back to the literature to look for some old papers and see what our chemical intuition would be telling us about phenylalanine and tyrosine and their propensity to be uh, phase separating, about their stickiness. And the first thing we looked at were were solvation free energies. And in the case of uh, well, these experimental data set for uh, side chain analogs, you can see that the tyrosine solvation free energy uh, is much more favorable than that of phenylalanine. So, in this case, uh, I think it's quite clear that you would expect a phenylalanine to be the stickier residue rather than the um, tyrosine. Or at least, that's what we thought. Another piece of evidence comes from uh, very elegant analysis carried out by Julio de Se and Christian Lindorff-Larsen and their co-workers, where they were looking at many hydrophobicity scales, uh, which were summarized in these stickiness parameter. And they looked at the distribution across the many hydrophobicity scales of this stickiness parameter for all the amino acids. And what they found, if we focus on the two residues uh, that we are studying is that again phenylalanine would be the more hydrophobic residue or the more sticky residue while tyrosine would appear at much lower values on average. So again pointing to the fact to the expectation that phenylalanine would be the stickier residue contrary to the experiments. Then we went into statistical contact matrices that have been have been from from experimental data sets of protein structures. You can go into the PDB retrieve lots of structures and look at the frequency of contacts for different types of pairs which would be captured by these NIJ and using a suitable references state you can convert this frequency of contact formation for different pairs into some pseudo energies. If you look at the most famous of these statistical contact matrices for phenylalanine and tyrosine you would again find that from this type of arguments you'd expect phenylalanine to be the stickier residue rather than what was finding the what was found in the experiment which is telling us that tyrosine is is much stickier than phenylalanine. There's also some information suggesting that tyrosine tyrosine interactions are stronger than those of phenylalanine phenylalanine. You can look at potentials of mean force like these calculated by Colipardos Rosana Colipardos team and you'll see that the tyrosine tyrosine PMF is deeper than that of phenylalanine phenylalanine so that's evidence pointing in the same direction as the experimental work on condensates. And also you can go into these even more enigmatic data sets on solubilities where you will find that in fact tyrosine appears over an order of magnitude below in solubility than any other residue or for that matter phenylalanine, again pointing to a greater stickiness of tyrosine. So, depending on what you look at, you'll find that either phenylalanine or tyrosine is a stickier residue based on naive expectations. So, we wanted to go a little bit beyond these naive ideas or our chemical intuition from other published data sets, and we thought that maybe what could capture these greater propensity of different amino acids to phase separate could be modeled using free energy calculations that are used extensively by our community, uh e- particularly when you're trying to study well, you're doing drug discovery studies and you want to look at whether one molecule or another molecule will be binding more effectively to a partner and by virtue of a thermodynamic cycle, you can do an alchemical transformation both in both the states in both the free and bound state and indirectly get the delta delta G binding and and you do this by invoking non-physical intermediates that are easy to simulate. So, we thought we would do this for condensates and specifically what we wanted to measure in our calculations was the different propensity of transfer of a phenylalanine peptide into a given biomolecular condensate or tyrosine being transferred into the biomolecular condensate. And we thought that maybe if we were getting that the transfer of tyrosine was more favorable than that of phenylalanine, then we would be capturing this greater greater propensity of um tyrosine to phase separate that had been found in the experiments. Uh again, we used one of the BioExcel tools to run these calculations. Specifically, we used uh the the very elegant uh PMX non-equilibrium alchemical uh switching method by uh Bert de Groot and Vidutas Gapsys and Mateus Delli and and their co-workers, uh where one typically requires uh two different uh simulations for uh both lambda states that one is considering. In our case, one will be phenylalanine and the other uh will be tyrosine. And after running these long equilibrium simulations, you will do uh many non-equilibrium switching runs uh both in the forward and reverse direction from lambda zero to lambda one and from lambda one to lambda zero. And by computing the distributions of work for the forward and reverse tran- non-equilibrium transitions and using uh non-equilibrium theorems, one is able to determine the uh delta G of interest. How does this look in the case of uh condensates? Well, in the case of condensates, we have to first set up a simulation box. What I'm showing here is the density across the simulation box, and you see that there's a dense region and a dilute region. So, here we have the phase separated system, and this red line in the center is the peptide that we are going to alchemically transform. This we did at lambda zero for phenylalanine peptide and lambda one for the uh tyrosine peptide. And then we run the switching runs. And this type of calculation can be done in the condensate, like in the case of this plot here or this other plot here, but it can also be performed in water. So, from those calculations, one will be able to obtain all the delta Gs that can then be combined to obtain the difference in transfer free energy between phenylalanine and and tyrosine for these model systems. And this is how our simulations look like. We have the condensate at the center. We have lots of water molecules in the dilute phase that sometimes you'll see that the peptides exchange with the dilute phase. The faint red in the center, that's our alchemically transformed amino acid residue which is at the middle of a slightly longer peptide than the ones that form the condensate. So, we did these very many times for multiple condensates. And what we obtained was that when the condensate was formed by glycine, serine, and tyrosine, the delta delta G transfer was negative, telling us that the transfer of tyrosine, the tyrosine bearing residue, was more favorable than the phenylalanine bearing residue. So, in good agreement with the experimental work. We replaced tyrosine in the condensate for phenylalanine to see whether the tyrosine in the condensate was responsible for this result, but we found that for a condensate formed by glycine, serine, and phenylalanine, the result was very much the same. So, we thought that probably it was the interactions that this tyrosine was making, which would be different from the ones that phenylalanine would be making. And we looked at the interaction patterns in this central region of the protein, so of the condensate we we run additional simulations of dense phases alone. Uh we looked at uh statistics of contact formation for tyrosine and phenylalanine. We found them to be essentially indistinguishable regardless of which was the residue that was part of the condensate. We looked at the interaction patterns and regardless of whether we were looking at pi stacking or sp2 pi interactions as a function of these angles and distances, we would find that essentially there there was not that much of a difference in terms of the interaction patterns that were determining these um result. Uh of course, there's a main difference uh that between phenylalanine and tyrosine and is that tyrosine can form hydrogen bonds, but also one has to notice that hydrogen bonds can also be formed outside, so it is ambiguous whether this is making a strong contribution in terms of the energetics for giving a preference to tyrosine relative to phenylalanine. And we thought that the condensates were in some way acting like uh a solvent that was contributing a medium uh and in this medium, uh which is highly hydrated and is capable of of forming hydrogen bonds, uh well, this was making being preferable for the for the tyrosine than it was for the phenylalanine and this is what's determining the greater uh propensity to to phase separate. But if condensates promote this trend, maybe different solvents would promote a different trend. So, we run these alchemical transformations and computed these delta delta D transfer into many other media um like acetone and methanol or ethanol, all of which can form uh hydrogen bonds, but have a lower dielectric uh constant and we found that these negative well this preference for tyrosine actually was reversed as we were increasing the size of the aliphatic side chain of the alcohols. So there seems to be a crossover between a tyrosine preferring and a phenylalanine preferring region and this is accentuated even more when you go into really a polar solvent where the delta delta D transfer favors clearly the the phenylalanine relative to relative to tyrosine. And by looking at the dependence of this delta delta D transfer with the with the dielectric constant we found that the dielectric constant seemed to track this crossover between the phenylalanine favoring region and the tyrosine preferring region. So after well obtaining these results my colleague Xavier Lopez devised this alternative way at looking at the same problem and proposed this a scheme for running quantum chemical calculations for a contact formation thinking that it would involve the well two amino acid residues I and J and that it could be split in two different steps. The first step would involve the transfer into a medium maybe the condensate or maybe a different solvent and then there would be the contact formation step and luckily with quantum chemical calculations we have methods that enable for us to calculate both the steps very accurately. So one of them would correspond to the transfer step that can be calculated with SMD very accurately. and additionally one could also try to obtain the interaction energies and combining all of this then what would be able to get the contact formation propensity into the medium that we're trying to to study. Now we did this types of calculations for many different arrangements of phenylalanine and phenylalanine and tyrosine and tyrosine including also their interactions with backbones that are going to be present in the biomolecular condensates and can be partners of these amino acid residues and um summarizing many many calculations we represented the difference in these contact between any combination of the of the tyrosines of phenylalanines with respect to our reference which would be the most stable of the of the phenylalanine pairs and what we found is that at intermediate die electrics we see again a crossover into a regime that favors tyrosine relative to phenylalanine. Also this is dependent on the type of solvent in which we are doing the transfer which in the is indicated by the different symbols and and depends on the configuration at hand specifically when the solvents are well don't have hydrogen bonding groups you always get that the that the well phenylalanine is is more likely in the case of alcohols or a special type of water where we tuned the dielectric constant you get into this region that makes tyrosine preferable. So we seem to have identified a crossover between these two regimes and we think that these regime of low dielectrics is actually what corresponds to the the course of folded proteins which are characterized by low dielectrics while this region of intermediate dielectrics is actually what corresponds to the tyrosine favoring region of biomolecular condensates. We've calculated these dielectric constants for the condensates. They seem to be in the range of 30 and 40. So, in this case we may be in the in the tyrosine favoring regime. So, after settling this we thought we could continue with our work and go into the other sticker strengths in those of arginine and lysine. Both positively charged amino acids that can form interactions with aromatic residues via cation pi pairs and again other people have looked into this before but we did this using our approach using our thermodynamic cycle, the alchemical switching transformations from PMFs for the transfer of either arginine or lysine, lysine here and arginine here into the biomolecular condensate. And the results in this case for this type for this pair of residues is different in all cases we obtain that the transfer is favorable for arginine again in in good agreement with experiments. There's no clear dependence with the dielectric constant and what was very interesting that came out from these uh, quantum chemical calculations that we performed also with support from a very talented undergraduate students, uh, Lydia Armentia from our lab, is that specifically the, uh, well, lysine seems to be in in the very, uh, well, low propensity to desolvate, uh, for lysine, uh, which is, uh, well, much more difficult to transfer into the condensate that arginine uh, in the case of arginine, the positive charge is delocalized and that favors, uh, the transfer into into the condensate. So, I will conclude here with these, uh, last results. Uh, I've tried to tell you about how these, uh, LCR-like amino acid mixtures that we've been putting together using these, uh, simple models seem to, uh, capture some of the interesting properties of intrinsically disordered proteins that are able to phase separate. Then I've told you about our work of on tyrosine and phenylalanine using, uh, alchemical transformations and DFT calculations where we identified a crossover in interaction strengths that we believe explains the different propensity to phase separate, but also reconcile many observations relative to the, uh, relative, uh, strengths of interactions inside the cores of folded proteins. And, uh, finally, I concluded with uh, just a few comments on our more recent work on arginine and lysine where the hierarchy is, uh, always, uh, the same regardless of the environment due to the, uh, large penalty for desolvation of lysine. Uh, this work has all been published in the references I'm showing here, so maybe there are some more, uh, more details uh, there, uh, if you're interested. and uh, with this I finalize. I would like to thank you for your attention and also again bioexcel for hosting the seminar. >> Thank you very much, David. And we have a couple of question in the chapter, so I will read it. So, the first one is how did you calculate the aggregation propensity? This was at roughly the middle of your presentation. >> Sure, sure. Yeah, so thank you for the question. Yeah, the aggregation propensity and I cannot exactly remember what the other parameter was. Yeah, let me just go back. These are obtained from the two measurements that I showed first. So, you can calculate the the change in the accessible surface area from the perfectly dilute regime and also you can calculate the maximum the maximum cluster size. So, from the cluster size you get the cluster what's called the clustering degree. And the aggregation propensity comes from how much you shrink across your simulation from the perfectly dilute state. So, in some cases you will see that there's a large decrease in the solvent accessible surface area. That is very much the case of the tyrosine which excludes all the water and then you get a large reduction in the accessible surface area. That results in a high aggregation propensity. For other systems, you get very much you stay at very much the same accessible surface area because they remain perfectly dilute and in that case you have a very low aggregation propensity. These are not parameters that I have defined Uh, myself. And if you go to the Biophysical Journal article, you'll see the origin for these definitions that were established by other people, not by me. >> Thank you. And then there is another question that it was was say thank you, really nice talk. Uh it was wondering that if your results were consistent across force field. And or there is one that perform better than another referred to force field. And then also go on how the water model can affect the real behavior of the condensate. If so, in your opinion. >> Yeah, so I guess I should start by saying that we haven't done a systematic study of different force fields in our work. I think this is actually something that may be extremely useful because I I guess that one would be able to get these calculations quite rapidly and and and across force field. So maybe it's a it's a good place for doing a force field validation exercise that we really haven't done. We have done some tests though and I guess I I started by by mentioning that we've used the our work primarily relatively old force fields. So much of our work has been done with FF99 SB12DN. Uh we tried the same systems with FF03 star with backbone corrections. It yielded very much the same results. But then we obtain a very different result in the case of the 99 SB disp force field which also comes with a different water model. Uh in this case, we wouldn't see phase separation for the ternary system. Uh this is something that was not too concerning for us because I guess some deficiencies in binding properties have been identified for these 1999 1999 SB disp force field. So, in that respect, maybe this this is a result that we could expect. Um we're doing now work with uh force fields optimized by Jain et al. precisely for the study of condensates. I would say there is going to be definitely dependence on the models that you are using and maybe the force fields that we have been using and with the water model, the tip 3p water model, they are more collapse-inducing. Uh our recent work and the work of many others on the small model systems seems to indicate that with a degree of variation, uh well, there is consistency across the board that the type of small peptide models that we're using, well, seem to be forming these types of condensates, but I would agree that this is probably something that requires careful validation against experiment. >> Thank you. So, we have another two question. So, the first one is where the simulation performed for GSF or G GFY amino acid mixtures? And then there is a follow-up, did you use any peptide? If so, what length? And would there be any limit to the length of the peptide that can be simulated in this way with this approach. I guess. >> Yeah, so I guess there's a number of different things there on the GSF GSY part. We did both. So [snorts] in the in the first paper in biophysical journal, we only reported like the ternary mixture with the tyrosine well as an as an ingredient of the ternary mixtures. In the more recent work from last year, we did both the GSF and the GSY and the reason was not so much that we were interested in the different mixture as such but on the possibility that the propensity we were finding for tyrosine was due to it having other tyrosines to interact with. So we wanted to swap the component of the mixture to see whether interaction patterns would be drastically affected with maybe much more pi stacking in the case of phenylalanine but that was not the case. So we did both GSF and GSY. Then I think there was another part of the question which related related to the size of the peptide. >> Well, I guess >> And then 80 amp used one was asking which length. >> Yes, so so I guess that we are using the dipeptide or the terminally terminally blocked residue monomers for the condensate but the alchemical transformation is performed on a pentapeptide with an residue that changes the residue that is alchemically transformed is the one that can either be phenylalanine or can be tyrosine. This is a peptide that has been this pentapeptide has been studied using NMR and and many different uh methods in the past. So, I guess it's a good test bed. It simply has a sequence GGXGG. So, it's just glycines and a host residue at the center or the guest rather residue at the at the center. And in all cases, the peptides are terminally blocked. So, both for the amino acid mixtures and for the peptide we are alchemically transforming transformed they are not zwitterionic. They are uh neutral terminally capped uh residues. And I think there was one one more thing, no, Alessandro? >> Yeah, so but you mute one. You try you alchemical transformation is involving one residue. >> Yes. Yes. Just one. >> Okay. Yeah. In in a pentapeptide? >> In a pentapeptide. >> Yeah. And then the last question was if there are any limit in the length of the peptide which for this type of simulation. >> Yeah. So, I guess uh well, in terms of limitations could come in in in in different ways. Uh I guess that one will be on based on computing power and probably our computing power will end up uh early. So, that's why we go small and and decide to to study like these uh uh little uh systems. Um when I was before in a different question claiming some robustness in the results regardless of force fields is because we're not the only ones doing this. Uh people have been looking at these peptide condensates for systems that I would say are up to some eight residues. Uh you should look at the work by David Botoyan, Daryl Joseph, and uh many other uh people who are looking at these types of condensate peptide uh models um as a proxy for the bigger systems. and we seem to all be uh getting consistent results. Tetrapeptides have been done a lot at the beginning of I talked I was showing simulations of full length IDRs for proteins of over 100 residues. So, depending on your computing power and the difficulties in setting up the simulations that I mentioned at the beginning, I guess you can simulate as far as you can. >> Yeah. Thank you. I I think we we still have two questions, but I take only one for reason of time and I will ask you to briefly answer because we are running out of time. Thank you. And so, they thank this the attendee thanks for the great talk. As you have nicely shown, thickness induced by tyrosine and phenylalanine is one of the reason of the formation of the condensate. And other is charged residue and more precisely charged patterning. What is in your opinion on the relationship between the two? Do you think that a thick patcher compete with charged patches or on the contrary they are complement or complete with one of with the other in the disorder in intrinsic disorder protein? >> Yeah, so that's a a great question and one that probably would be better answered by some people who've looked at the sequences in more detail. I think that an important point is that there's different types of sequences that are involved in phase separation. What we've been looking at are aromatic driven and also cation-pi driven. So, in our case, when we're looking at the charged positively charged residues, we're looking at the interaction of these positively charged with a condensate that has lots of aromatics. In this case, would be complementarity, not competition. But, I guess I should also mention that there's also biomolecular coacervates that are induced by charged positively charged with negatively charged amino acid residues. And in that case, you don't really need the aromatics. So, I guess condensates come in in all sorts of shapes and and forms. And I'm not sure there's much of a competition between both types of possibilities. I think they're just different. >> Okay, thank you very much. I will ask kindly the last question to if the attendees can post it on the forum. And then we I thank you a lot to David. I just do a small announcement and then we close this session. So, I want to just to tell you about the next webinar and the last for this spring season. And that is the 26th of May where Kerstin Lindorff-Larsen will speak about conformation ensemble of intrinsically disordered region and protein. And last not least, we have postponed the abstract deadline of our weekly bi-yearly conference that will be the 31st of May. So, there are the deadline of abstract the conference will be in September. And we'll cover bio uh sorry, multi-scale modeling, molecular dynamics, free energy calculation, integrative modeling, drug design, and AI. So, please come so we all meet in person. It will be great. So, don't miss the deadline. Thank you very much again to everybody, and we see in two weeks. Bye-bye.