Submind YouTube summaries
Thumbnail for Bioexcel webinar #98: Conformational ensembles of intrinsically disordered regions and proteins

Bioexcel webinar #98: Conformational ensembles of intrinsically disordered regions and proteins

Watch on YouTube

Video summary

Kerstin Lindorff Larsen from the University of Copenhagen introduced computational methods designed specifically for studying conformational ensembles within intrinsically disordered proteins (IDPs) and regions, emphasizing that while only about 5% of human proteins are fully disordered, most contain a mixture of ordered and disordered segments. Traditional modeling approaches relying on large structural databases like the PDB or multiple sequence alignments often fail to capture this biological context because IDPs lack well-defined native states and evolve differently than folded proteins. To overcome these limitations, researchers developed coarse-grained simulation models parameterized directly from experimental data such as NMR paramagnetic relaxation enhancement and small-angle X-ray scattering (SAXS), utilizing Bayesian optimization to fit force field parameters without depending on predefined ensembles or synthetic HP-models initially used in early iterations. The resulting frameworks, including the Cabadas model which incorporates non-electrostatic interactions and hydrophobicity priors derived from literature scales, successfully predict chain compaction based directly on sequence features like charge distribution and blockiness. This capability allows for the rapid analysis of thousands to millions of sequences without extensive simulation time, effectively bridging the gap between folded and disordered states across a continuum defined by chemical shifts and SAXS data. A key application involves integrating these models with tools like AlphaFold2 or ColabFold; while standard AlphaFold excels at predicting specific structures for folded domains but struggles to generate the necessary structural diversity required for IDPs, newer approaches freeze well-folded regions using harmonic constraints while allowing disordered segments to be simulated dynamically via direct experimental data fitting. Further analysis highlights that current models rely exclusively on solution-based experimental data rather than crystallographic information and balance electrostatic interactions with hydrophobicity through screened Coulomb potentials featuring fractional charges near pKa values, despite some limitations regarding salt dependency or varying dielectric constants. Although all-atom molecular dynamics offer detailed local structural insights, they often struggle with convergence for long IDPs exceeding 40 residues, making these efficient coarse-grained models superior for studying global properties and multi-domain proteins where disordered regions behave distinctly in full-length contexts compared to isolated loops. Future developments aim to expand benchmarks across mixed-order systems, integrate post-translational modifications and molecular crowding effects, and enhance nucleic acid modeling by adding harmonic constraints for duplex structures beyond current bead-based representations lacking sequence specificity. The presentation concludes with a comprehensive evaluation of generative models like Pep-Zero against established benchmarks such as PeptoneBench, demonstrating that these new tools outperform traditional methods specifically on IDPs while maintaining accuracy on folded proteins when combined appropriately. By addressing the critical need for structural diversity and avoiding the loss functions inherent in standard AlphaFold approaches, this research provides a robust pathway to understand how intrinsically disordered regions function through chemistry-specific polymer properties and short linear motifs. The session also acknowledged the development team's contributions, invited further discussion via an online forum, and announced upcoming deadlines for the BioExcel conference events scheduled in Brno, underscoring the collaborative effort required to advance computational biology in this complex field.
Read the full video transcript
Welcome everybody. Today we have the BioExcel webinar number 98. And with us is Kerstin Lindorff Larsen from the University of Copenhagen. And he will speak about conformational ensembles of intrinsically disordered region and protein. I'm Alessandra Villa that together with Otto Anderson and Richard Norma we are hosting this webinar for BioExcel Center of Excellence for computational biomolecular research. This webinar is being recorded. I wanted you are aware of that. During the webinar you could ask question using [snorts] the Zoom application that you find at the top bottom of your application. Depends on the operating system you might see one of those three symbols. Please click it, type your question, and at the end of the webinar we will read the question for you and Kerstin will answer. After the webinar maybe you still have some question. You are most welcome to join us at the BioExcel ask.bioexcel.eu forum. And there will be a post for the webinar of Kerstin and then you can type your question there. Kerstin will be able in the following week to answer for the question. Something about Kerstin. Kerstin is a professor of computational protein biophysics at the DeLund Strom Center for protein science. He's director also of the PRISM protein interaction and stability in medicine and genomics center. And the member of the Royal Danish Academy of Science and Letters. He got a Danish Independent Research Council Young Researcher in 20 in 2006. He was co-recipient in 2009 of the Golden Bell Prize. Congratulations. And in 2025, the Novo Nordisk Foundation Prize for natural science teachers at the university. Currently, his research focus on developing applying computational methods for studying studying the structure and the dynamics of protein and integration of biophysics and genomics genomics research. And today we'll speak about intrinsically disordered region and protein. Okay. I stop sharing and welcome. >> Thank you very much, Alessandra, and everyone else who's online for joining. Uh, it's my pleasure to be here today to talk about some of our work over the years on developing computational approaches to study intrinsically disordered proteins. So, I have put together I let me just minimize this thing before I I've put together a series of slides that talks about our work on how one can determine conformational ensembles of intrinsically disordered proteins and and regions and focusing a little bit more on how we develop the approaches and then with with some pointers along the way also on how one one one one can use them. Um so I'm not going to give a huge amount of of biological background, but nonetheless just want to remind everyone that proteins with disordered regions are pervasive. And so again, if I can manage to get rid of things here. Here's some some statistics that that that other people have collected and and you know, I've annotated it here with sort of three uh sort of types of proteins. Uh these are numbers from the human proteome. Um so if we sort of define proteins that are sort of mostly folded as proteins that have you know, at most I think 5% of disorder, that turns out to be about uh 40% of human proteins. On the other hand, about 5% of human proteins basically do not have uh well-defined folded domains, which means that the majority of proteins in the human proteome contain mixtures of ordered and disordered regions. And so if the toolbox you have available only works for fully folded proteins or only works for fully disordered proteins, you're basically losing out of studying most of human proteins. And so what I will talk about is how we are building up computational models that allow us to study this continuum of protein order and disorder. Another way of looking at the same numbers is if you look at the amino acid level, about 30% of the amino acid residues in the human proteome are what we would say uh intrinsically disordered. These proteins do a wide range of things, and I put in this slide mostly as an advertisement for this wonderful review that Alex Holehouse and Peter wrote a couple of years ago about protein disorder ranging from the biological function to the molecular biophysics. And so, if you're interested in these proteins, you should definitely read this review as well as other nice reviews from over the years. And of course, the the the original literature also. So, proteins with disordered regions are involved in basically all biological processes ranging from cell signaling to to transport to you know, transcriptional regulation, etc. etc. And the way that they do this is via a combination of different types of chemistries. And some of these chemistries are Well, these rules are somewhat different from the rules that we're used to think about if you think about folded proteins. And so, one way I think about it is that there are these disordered regions that can either be at the ends of proteins or between folded domains that have chemistry-specific polymer properties, which means that it's not so much the specific amino acid sequence, but it's more the overall properties of these these regions that drive some part of the function. And then there are other parts of function, for example, in so-called short linear motifs that are more sequence-specific that often govern more more specific interactions. This is at the very high level, but of course, in in in in many individual cases, we also know exactly, you know, how these different properties, you know, play together and how they determine biological function. One of the difficulties in moving forward in this field has been that the rules and the techniques that are used to sort of read and write sequences are somewhat different than the rules that we are sort of getting used to for for reading and writing protein sequences for folded domains. And so, I think it's sometimes constructive to contrast the situation of disordered proteins from the situation for folded proteins. And so if we for example think about the data sources that have been been used to develop structure prediction models or protein design models for folded proteins, they consist of two sources of information. We have the wonderful protein data bank with a few hundred thousand examples of sequence structure pairs, and then we have you know large large sequence databases, UniProt, etc. that have lots of sequences, and then these sequences we can then align in so-called multiple sequence alignments, and then we can train deep learning models that relate sequence alignments to structure, either sort of to predict structure or to generate sequences. Or we can use other models that maybe don't explicitly rely on alignments, but often implicitly rely on the fact that sequences are conserved in very special ways. For disordered proteins, we don't have these. We have the protein ensemble database that's small and heterogeneous. We have lots and lots of sequences, but it turns out that the way that these develop over evolutionary time scales are quite different from the constraints that operate on folded regions, and so they can change the sequences in very different ways, which means that multiple sequence alignments sometimes give you useful information, but other times the alignments really do not tell you the picture. And so we've been thinking about not just our group, but many other groups have been thinking about how can we learn sequence ensemble or sequence property relationship for disordered proteins in other ways, and this is what I'll talk about today. Some of these thoughts are summarized in this recent review that we wrote, and so this will be sort of the outline of at least part of the presentation. And the the major point is that if we don't have huge amount of structural data, but we still care about understanding structural properties, how can we learn those from the data that we have? And so, what I'll talk about is how we can take experimental data mostly from NMR and small angle x-ray scattering data, and use that to derive physical models, in this case coarse-grained simulation models that we will parameterize based on experimental data, and I'll tell you about how we do this. And then how we can validate these models that then also turn out to be relatively fast. And then we can use these to run simulations and generate conformational ensembles of thousands of of sequences that we can then use as input to train machine learning models if that's what we want to do. Or we can just use the simulations to give us some structural insight. Or we can use them to interpret experimental data. I will not really talk about that, but but but I will put in a a pointer at the very end. We can also take these models and sort of invert them either with the simulation tools or with intermediate machine learning models to do sequence design, and again I'm not going to talk about it, but it's described in this review as well as another review we recently wrote. So, the goal is how can we learn when we don't have a huge amount of experimental data? And I decided for this talk to go back a few years sort of to explain you how we got into this and some of the early work that we did um and how this sort of influenced our work later on. So, this is some work that was done by a fantastic master student I had in the lab now about 20 years ago when we started working on this problem. So, previously I've been working on disordered proteins and how we can interpret experimental data on these, but back at the time we thought maybe we can develop prediction tools that we can then sort of use to study when we don't have experimental data. And so, Anders together with a collaborator, Jasper Borg, uh came up with a procedure of how we can parameterize a coarse-grained model for what we call structure prediction of disordered proteins. Now, this was done in the context of Monte Carlo simulations, what I'll talk more about later will be, you know, molecular dynamics, but but but but the basically the models are are very similar. So, we want to have a model that is, you know, as well as we can, be predictive. And the work is published again some years ago, but I'll give you some few snippets just to to give you, you know, an idea of of how we we think about these problems and how we've been sort of developing our thinking. So, if you have a model like this, you need to parameterize it. And one of the ways that we can parameterize coarse-grained models is what we call sort of bottom-up parameterization. We have a more fine-grained model that we can then represent. This we could do for the backbone potential in a coarse-grained model. And in in for those of you who don't think about C-alpha level models, the C-alpha based models have a backbone that you can have a backbone potential for. In this case, there will be an angle and a torsion that determines the local geometry. We could actually generate reasonable backbone structures using a series of tools. Back then, there were these statistical coil models that operate at the full atom backbone level. And then we can project those down into this angle and the dihedral angle uh so, the angle and the dihedral angle. And then we can parameterize the coarse-grained model using these mixture models here that we can then combine to create what is basically the equivalent of the Ramachandran map, but for a C-alpha chain. So, this is a relatively simple way of generating a a force field is to start with a more fine-grained model and sort of, you know, parameterize it uh into into the coarse-grained model. But the more important part of the model is really the long-range or the non-bonded parts. And we decided that this would be better to do by tackling, you know, some some some some some experimental data. And again, these are some old methods and you know, these days we don't do these kind of things anymore, but just to remind you that there are a bunch of ways of creating statistical models for force fields, but very often they rely on having sort of target structures. We might want to, for example, you know, minimize, you know, or come up with a force field that you know, optimizes the Z-score so that we can distinguish the native state from other states. Or there's this fantastic old work on funnel sculpting where you maybe want to correlate the RMSD to the native state with the energy. And you can optimize the force fields against these things. And there's a bunch of different ways of optimizing, you know, force fields that, you know, stabilize a well-defined native state relative to some other states. But all of these models generally relied on some well-defined target structures. And of course, for these folded proteins, we don't have a native state, or rather native state is a broad conformational ensemble. So, we thought that these models were not really the right way of thinking about it. We and many others have been thinking about how one we can generate conformational ensembles. I'm here showing you two conformational ensembles of a protein. It's not so important what what it is. What is important is that these are two conformational ensembles, one on the left and one on the right, and they're both derived from the same experimental data. In this case, from NMR experiments. And you can see that they're quite different. So, it turns out that there's actually quite many different ways of generating conformational ensembles that are compatible with a single set of experimental data. And so, we were not super comfortable on targeting the experimental ensembles. So, we could generate experimental ensembles based on the data, but given this uncertainty of the ensembles, we thought, well, maybe this is not the right approach. So, instead, what we came up with was an approach of optimizing the force fields directly against the experimental data. And so, in this paper that that Anders published, uh really maybe the most important part of the paper is is described in the supplementary information, and it's a Bayesian approach where you formulate the learning of the force field parameters as an optimization problem given some set of experimental data. And so, the equation here at the top is basically base formula that says, what is the probability of some force field parameters epsilon given the data can be formulated as you know, relating to the probability of serving some calculated data or some data given the force field parameters. And this is something that we can then calculate because if we have some guess of what what the force field would look like. So, if you have some hypothesis of what the force field would look like, we have some hypothesis of what these epsilons would be, we can then run simulations, and then we can calculate whatever we might observe in an experiment, and then we can compare that to the experiment that we have. And so, this is basically what we derived. The equation below is the same as the above under some simplifying assumptions and taking some logarithms. And basically, the idea is that we have sort of a chi-squared term that measures the deviation between what we measure experimentally and what we predict from a simulation, but where the prediction from the simulation is dependent on the force field. And so, we can optimize this log-likelihood function basically by minimizing this. So, that's the basic idea. Of course, it's it's it's it's uh conceptually uh simple to write down. The difficulty is how you implement this in practice. And so, this is what what Anders worked on. Before I get to that, I'll just, you know, showcase you one of the types of the NMR experiments that we can do. This is paramagnetic relaxation enhancement NMR data. This will just be important because that's the few points of data that I'll show you. If we put in a spin label at a site, we can see broadening of NMRs throughout the chain in a distance-dependent fashion. And again, under some simplifying assumptions, there's a simple relationship between the average distance between the spin label and the protons or the distributions of distances that gives you this paramagnetic relaxation enhancement term that then influences these insensitive ratios. And although this is not the perfect measure, this is what we had available when we did this work again a long time ago. So, if we have a conformational ensemble, we can calculate the PIE and we can compare it to experiment. So, this is the experiment that I'll I'll I'll talk about. And so, we came up with this iterative optimization scheme where we have some experimental data. I'll show you how it works with the synthetic data. And then, I'll show you how we can optimize against this and show that the method works. And so, the idea is that we start with some parameters that we, you know, that we, um, that we guessed. They could, for example, see we set all the attractions to zero. We run a simulation to generate a conformational ensemble with that initial parameter set. We calculate the experimental data as, you know, observables. And then, the key trick is that then we can optimize the parameters so that we minimize the agreement or maximize the agreement or minimize the difference between what we predict from the simulation and what we measure experimentally and use that as a guiding function for the force field parameters. And then, we can iterate over and over again. And then, the whole trick is how do you do this without having to run so many simulations to test all possible force fields. That's really the difficult part. So, again, to test that the model work, honest came up with a nice synthetic data set. He said, "Well, the simplest way we can represent a protein is sort of what we might now call a stickers and spaces model." Or in the protein folding field, uh for for polymers, this was the the HP model. Um but it's basically a model where there's two types of residues. There's hydrophobic residues that interact favorably with each others, and then there's polar residues that don't interact favorably with each other or with the hydrophobic residues. And so, we created a synthetic uh data set. These are the black lines here that if we took this protein called ACBP and simulated it, we would measure this data set. So, this is now our synthetic data set. And then we say, "Well, we don't really know this one, so we'll start by setting all the interactions to zero." And the red lines is what you get if you run a finite-length simulation, you know, so with some noise um using a model where all the interactions are zero. And so, we ask the question, "Can we recover this HP model um from this experimental data?" And so, the approach is again relatively similar. We run a simple We run a simulation. We calculate the expe- experimental observables as these weighted uh averages here. Now, these weights here would typically just be uniform weights because we're doing in this case Monte Carlo uh sampling. And so, we can just run simulation. We can calculate these numbers. We can compare with the experiment, so we can calculate a chi-squared term um for what these these these these would look like. And now, the trick here is that we can now actually calculate the derivative of this chi-squared term with respect to the force field parameters. In this model, the force field parameters are called Q. So, we can calculate the derivative of this with respect to the force field parameters to do a gradient optimization where we minimize this chi-squared term by optimizing these Qs. And of course, this is difficult because we have a big conformational ensemble, so we need to propagate the changes in the force field through the conformational ensemble, but then Boltzmann comes to our rescue, so we can basically, at least for local changes, we can calculate how the energy of each of the conformations would change if we update the force field. So, we can propagate the changes in the Qs back through the ensemble to optimize this. And so, the gradients back then, we would calculate using finite differences. Then, of course, these days one can do this with automated differentiation tools, uh but back then we did it, you know, using this simple approach. And then you can then basically optimize the force field at least locally, and then you have to resample using the optimized parameters, and then you iterate over and over again. So, how does that look in practice? Here's one example. Here, the black line is as before, this is the synthetic experimental data. The green line is what you get from a force field that is bad. But you can then take this green line, and you can re-weight it by optimizing these force fields, and you can do that at least to some extent, and you get this noisy red line, where we optimize the force field parameters to get closer and closer to the experiments without doing any resampling. This gets noisy because we are stuck with the original ensemble, but once we are close enough, and that's described in the paper what close enough means, we can then rerun a simulation, we get this blue, and you can see this blue line is now closer to the experiment than the green line, meaning that we optimized the force field in one single step. And then we can basically do this over and over again. This is what's shown here on the right. We start here with the red, and then we move continuously closer and closer to the black line, which is the experiment. And we can see over here what happens with the force field. We start off with all the force field parameters being zero, and very quickly this optimization model learns that this HH interaction should be favorable with a minus one energy, and the other one should stay around zero. So, this method sort of works in the sense that we can recover a synthetic force field that we predefined. This is a simple model with, you know, three parameters. We also showed that we can recover a 20-parameter model, and we also showed how one can use it for experimental data. So, this is basically a approach of learning force field parameters directly from experimental data without having to define conformational ensembles. And so, one might have think, well, that was the, you know, the end of it, but of course, for us this was really the end of it, but only because we sort of didn't really continue this work. And one of the reasons was that at the time there was really not very much of this kind of experimental data available. And while there was some experimental available, this was often done on a very different set of conditions. Lots of the data were for unfolded proteins that were denatured in denaturants and acid, and there was not a lot of systematically collected data for intrinsically disordered proteins, at least not that we were aware of. Um but over the years, this has certainly changed. And so, over the years, lots of other people started to do similar things, you know, to develop coarse-grained models either from top-down or bottom-up, and validate these using experimental data. And then, some years ago, we decided to, well, maybe we can revive uh this idea of learning force fields directly from experimental data, but using newer tools, faster models, and most importantly, a larger set of experimental data. One of the reasons why we and others also uh got back into this question was that there was a research interest in studying just not just individual disordered proteins, but also interactions between disordered proteins. And so this figure is sort of meant to illustrate you that, you know, if we learn the rules of intramolecular interactions, so between different amino acids within the single chain, we also learn something about the intermolecular interactions between chains. And this turns out to be important if you want to study cluster formation or for example phase separation. And although I'll not talk about this, I'll just show you here some some some some experimental and computational data from a bunch of groups. These are sort of some key papers, but there are several others from over the years and also prior work in the in the polymer field. And then each of these plots, the stuff on the x-axis is related to how compact a disordered protein is, either by measured experimentally or from simulations of various types or competing or theoretical models. And on the y-axis are various things related to the protein's tendency to self-associate and for example undergo phase separation. And these correlations show that there is this intrinsic correlation. And so we use that as as sort of motivation to say that if we can learn a force field by looking at single chain properties, so this is the stuff on the x-axis, maybe it tells us also something about these proteins' tendency to undergo phase separation. And there's of course other people had had thought about this and sort of we just piggyback on on on these ideas. And so the idea was sort of picked up again initially by Ramon and Tea and then later joined by several other people, most notably Julio, and but Fanny, Key, Soren, Ariane, Anaida and and and several others have contributed over the years to sort of redeveloping this data-driven coarse-grained modeling. And so the model I'll talk about briefly is called Cabadas. It's a copy of other people's models. It has some similarities to our our previous models, but also some some additions. So for example, there's now molecular dynamics, there's now charge interactions, and and a few other differences. Um but but the key key parameter in this is still these non non-local interactions between amino acids, in particular these non-electrostatic interactions that are parameterized by this parameter that we we call lambda that says something about whether the amino acids like to interact with each other or whether they more like to interact with their solvent that we're not representing explicitly. And so we call this sometimes a hydrophobicity parameter or a stickiness parameter. So again, we took out the the the the model from from from Anders, we we changed it in a few different ways, and actually in some sense simplified the parameter learning um in a way that turned out to be a little bit faster, but the basic idea is the same. We write down this log-likelihood function that depends on the lambdas. We have two terms here that measure the deviation to experiments, and then one term over here that we ignored in the previous work. This is the prior in the Bayesian formulation, and then we basically optimize or minimize this function that sort of, you know, maximize the agreement with experiments, and, you know, minimizes the deviation from the prior. And again, it's this kind of iterative, you know, sampling and and and re-weighting kind of approach. So for the experimental data, we went into the literature, we found experimental data for initially about 50 proteins. Now we are in more than 100 proteins, mostly from small angle X-ray scattering data measured by wonderful collaborators and non-collaborators over the world who measures data, puts them in databases, make them available so that we and others can can use it. So thank you very much to all the people that measure measure data. That's fantastic. As well as NMR data again that we've collected from the literature or people have, you know, dug out of old folders and and so on and so forth. So we have a a medium-size data set that allows us to learn these parameters. Now, this updated version we also have a prior and so our prior was to go into the literature and look at hydrophobicity scales. And so we we took about 50 or 70 hydrophobicity scales, you know, normalized them and say, well, you know, maybe they don't agree all of them, but most of them think that aspartate is not so sticky and tryptophan is more sticky. So, at least in the absence of experimental data, this would be our prior expectation. And so we used that sort of to regularize the optimization also. And so this in particular plays a role for for for amino acid residues that are are less common across these 100 proteins or so that that we are now optimizing on. So, the optimization looks a little bit, you know, look pretty similar to what I showed you previously. I'll just play you a movie that sort of shows how it operates. We I know, we run a simulation with a force field and we're slowly optimizing it so that we maximize the agreement in this case, for example, with the radius of observation or over here the NMR data. Um but now doing it not just on a single protein, but you know, simultaneously on in this case about 50 protein and later on, you know, about 100 different proteins in in total. And so in total, we can optimize force fields by maximizing agreement with experimental data on a broad range of proteins as well as this prior that comes from the hydrophobicity scales. And so this is sort of roughly what the final stickiness scale looks like and it has some similarities and some differences to the original model. We call the model Cavitas, this is the acronym, it's not so important. We we've we validated it quite broadly as I'll show you in a second. This is sort of the agreement with the training data set. You know, this has a very high correlation mostly because it's pretty easy to predict that long proteins have a greater RG than small proteins. But also if you look, for example, of variants of the low complexity of A1 from from mostly 10 nanometers lab, we can see a very high correlation between the measured and the calculated radius of gyration. This is now within the training data. So at least this tells us that we can recapitulate these things. But we can also look at independent data that other people have have collected or we have collected over the years. Sex data, fret data or or sex data you know on proteins that all have the same length but where maybe the amino acids have been moved around. And again this is described in published literature. I'm not going to go through it. Just to say we collect data from the literature and again you know big acknowledgement to all the people who measure high quality data and make it available. And then we can use this in this case to benchmark the models. When we then look at this scale, we can see that this scale sort of makes sense in the sense that it looks like what people had suggested by looking at experimental data for disordered proteins that there are some residue types that like to interact more with each other than others. We've just published a paper where we sort of try to make this systematic also by saying that the radius you know trades we get get get rid of the problem that the different amino acids have slightly different radii. But we can then also go back and look at what kind of hydrophobicity scales to these to this scale correlate with. And what we for example we find that this you know scales that are down here, these are are sort of our new scales, they are similar to some experimental scales for example this Jure scale that others have been using to study IDPs in simulation that is parametrized again on sort of disordered proteins but is for example quite different from the commonly used Kyte-Doolittle scale and because some of the residues that are sort of sticky or or or interaction prone in the context of disordered proteins are not interaction prone in the context of of IDPs. And so, this scale, you know, in many ways recapitulate what other people had found in sort of individual cases, but maybe makes it also systematic across all different 20 amino acid types, learning it directly from the experimental data. So, what can we do with such a model? I'll show you uh you know, a few few snippets of what one can do. So, we can, for example, now start to explore sequence-structure relationships. So, how do we go from sequence to structural properties? In one example, Julio and Anna ran simulations of about 28,000 IDRs that we extracted from the human proteome, cutting out the IDRs of of of full-length proteins, and ran simulations of all of them, put them in a database that you can analyze. So, I'm not going to go through the details. This is all published. Just to say that we can sort of conceptually group them into things that have different levels of compaction independently of chain length. You can quantify this in different ways. One way is to take this apparent internal distance scaling exponent, that when it's low, it means that the chain is compact, and when it when it's high, it means that the chain is more expanded, but that is more independent of the number of amino acids. And we can then start to say, what are the sequence properties that determine this? And we find that these again recapitulate what others had found in individual cases, and sometimes also systematically. So, that there is this interplay between having a high or low fraction of charged residues, a high or low hydrophobicity, and if you have many charged residues, whether they are well-mixed or blocky along the sequence. And so, we could put all of this together in a single coherent framework that sort of combines these features. Again, not features that we develop, but put them into a single framework that say, what is the interplay between these across proteins that are natural sequences that determine whether they are highly expanded or highly compact. And basically rules are pretty simple. If you are highly charged and not charged blocky, you are typically expanded. If you are medium, you know, if you hydrophobic and not so charged, you are typically more compact. But if you are highly charged, but then charged blocky and a little bit less hydrophobic, you can then become even more compact by forming long-range interactions that are driven by charge-charge interactions in these kind of more blocky sequences. These rules are actually relatively simple, so you can develop very, very, very simple machine learning model to predict this number from sequence. This is support vector regression, so this is probably some of the simplest models that you could develop, but it turns out to work, you know, extremely well just for predicting conformational properties such as chain compaction directly from sequence. So, we don't even need to run simulations. We can just predict from sequence, and this is available for anyone to predict. This is convenient because that means that we can scale up instead of looking at 20,000, we can look at a million sequences, and we can ask questions. For example, are these properties conserved across evolution? And it turns out that the compaction of an ortholog is typically strongly correlated to the compaction of a human IDR, not always, but at least on average. So, this, you know, was is maybe not surprising. And of course, it depends a little bit on how you find orthologs. That's a whole separate talk. Um but it turns out that others had found that while this is true often, it's not always true. So, here's a really nice paper. This is not our work, that showed that there was this intrinsic correlation between the sort of per amino acid uh compaction or in this case specifically actually the end-to-end distance and and the length so that there is some compensatory changes in the sequences over evolution so that if the IDR becomes very long, it becomes a little bit more compact per residue. And so we can now start to see can we find equivalent cases like this where it appears that the global chain dimensions are more conserved than for example the length of the amino acid chain. And so again, this is published so I'm not going to go through it just to say that there are several cases where it appears that over evolutionary timescales there is a correlated change between the length of an IDR and sort of the sequence properties that drive compaction so that there's this negative correlation so as the chain gets longer across evolution or shorter for that matter, then the sequences adopt to preserve the physical dimensions for example the radius of gyration or the end-to-end distance more than for example the average amino acid composition suggesting that at least in some of these cases the physical dimensions might be under evolutionary pressure like in the paper that I just showed you previously where they demonstrated you know this has strong functional roles. So I started by saying that only 5% of human proteins are fully disordered so of course it's important to put this disorder back into context and we've been pushing the model to work for multi-domain proteins, nucleic acids, we're interested in crowding, PTMs and other things. And I won't have time to talk about all of this but I'll just show you briefly some work we've been doing on multi-domain proteins. We developed a model called Calvados 3 where we keep the folded domains rigid by a harmonic network constraint how you might do it in other coarse-grained models also. The model predicts radius of gyration and also phase separation relatively accurately, so it doesn't sample the dynamics of these folded domains particularly accurately, but it keeps the structure together as you would, you know, in in in other coarse-grained models. One limitation of this is that we need manually to define where the folded domains are, and this can be a little bit tedious if you want to scale up. So, obviously we came up with an alternative approach of doing this that we call AlphaFold ColabFold, there's a preprint out um that you can have a look at. And basically, we replaced this manual step, or rather that CERN developed this manual step, where he basically combined AlphaFold structure prediction together with PLDDT score and the AlphaFold PAE matrices to come up with a way of automatically constraining the folded domains and letting the disordered regions be flexible and governed purely by ColabFold. Then the folded domains are kept together with sort of a go-type structure-based model. And so, we can now easily take a sequence, run an AlphaFold prediction, and then input that into the AlphaFold ColabFold framework, and now generate conformational ensembles of a protein with both folded regions and disordered regions without having to do any manual assignment of the folded state. This model is as accurate as the manual model, so that's great, but it's automated so that CERN could easily run simulations of in this case 12,000 full-length human proteins that are, you know, found in the cytoplasm or the nucleus. And so, we put them in a database that you can go analyze. Just to give you one example of the kind of analysis we can do, we can now compare how the disordered regions behave when the context of the full-length protein versus in the context of, you know, just being isolated. And we're sort of mining this data to try to figure out what are the context rules. For example, we find some ideas that are much more compact in the context of the full-length protein than in when they're isolated. These are not big discoveries. This is what we commonly call loops. So, these are basically regions that are constrained because they're sitting either within a folded domain or between folded domains that are tied together. So, this is not a big discovery, but they fall out very naturally from this kind of analysis. But, there's lots of other variation that that is interesting and hopefully, you know, we and others will be using our data. This is our data for transcription factors. And in the final few minutes, I'll just say that, you know, for these kind of models and you other models out there, we really need better and broader benchmarks that operate not just on the disordered regions that I showed you we have the SAXS and the NMR data, or the folded regions where of course we have, you know, the wonderful PDB, but also of things in between. And so, a part of this other paper that developed a generative model either for IDPs or for folding protein, this is work driven mostly by people at the company Peptone together with NVIDIA, and we also played a minor role in in this work. So, I'm not going to talk about this model, but as part of this work, one of the key steps in this paper was also to put together a benchmark that operates across this order-disorder continuum. The model is called PeptoneBench. It has several hundred proteins that have been, you know, mined from the two databases, the chemical shift database BMRB and the SAXS database SASDB, automatically extracted so that we can now take structure models for proteins with mixed order and disorder and benchmark both their local structure by comparing to chemical shift and global structure by comparing, for example, to small-angle X-ray scattering data. So, we can now calculate chemical the and compare to experiments, and then put it on this RMS E axis, so when low means good and high means bad. But now, each of these points is an individual protein or individual data set that sort of spans this order to disorder continuum. And this means now that we have generative models or simulation models, we can now benchmark these at scale. So, here's just some example what this might look like. The top row is for chemical shift, the bottom row is for small angle X-ray scattering data. If you take AlphaFold just as a simple baseline for structure prediction, a folded protein, AlphaFold, not surprisingly, does amazingly well for folded proteins, but of course doesn't really predict chemical shift of disordered proteins because that's not what it was developed to do. That's true also for SAXS data. If you take this IDPO model, that's IDP specific, it predicts SAXS data very well for disordered proteins, but not so well for ordered proteins. Whereas this Pep-Zero model, for example, operates relatively well across the order-disorder continuum. So, we can quantify how well something works by basically calculating the area under this curve from the ordered regions to the disordered regions. And here's some benchmarks on different models. There's AlphaFold and both, there's BioLuminate Pep-Zero, and you can see, for example, over here for chemical shift over here for SAXS, that these more broadly applicable generative models do quite a bit better across the entire order-disorder continuum. So, whereas AlphaFold does fantastically for folded proteins because that's what it's meant to do, it doesn't do so well over here. And so, we said, "How well does our model do on this?" And we were very pleased to see that by basic copying structure of the folded domains from AlphaFold in AlphaFold ColabFold, but then letting ColabFold deal with all of the disorder, we actually created a model that for global properties, as measured by SAXS, actually does really well across the entire disorder continuum. So, it even does better than this bio immune and peptide model and these in this kind of benchmarks. So, this model, which is very simple, it has 20 parameters, um basically that interaction strength for each amino acid type, allows us to predict conformational ensembles of proteins with mixed order and disorder that is at least as accurate as these more complicated models. Of course, this model doesn't do everything that these models do, but it's very efficient and and and and and and and and as you can see, pretty accurate at capturing these conformational changes. You can run all of this by going to our GitHub page, either download the Calvados package and there's sort of a a manual here or a bunch of examples. You can also run it through Colab. Um there's the AlphaFold Calvados where we put in all the conformational ensembles that you can just download. If you know the UniProt ID, you can click through and generate our simulations. Or down here at the end, you can also run it yourself with a Colab certificate. And so, with that, I'm at the end. We can learn force field parameters uh by targeting experimental data. Um we can combine this with traditional machine learning models such as, for example, AlphaFold protein language model to build in specific interactions and we can run these simulations at scale and then we can generate training data for predicting properties from sequence. I showed you one example of this. I didn't talk about how we can use it for generative models. Um or we can invert these models again for sequence design. And again, I just want to say that we need better and broader benchmarks that are based on experimental data. I showed you one data set. I think it's very good, but we need more than this. I I think I mentioned the people along the way and that did most of this work, um but I'll just show you their faces here. I'll show their faces here at the end um for the people who developed first the early version of Calvados and now the more recent versions, as well as the people that did all the applications that I talked about. With that, I'll stop and happily answer any questions. >> Thank you very much, Kerstin. It was a great presentation. We We got also a lot of question in the chapter, yeah. So, I will So, some of them are tied to a exactly moment in your presentation. So, I don't know if they are understandable, but I will try to to be to go to jump from one to the other. And so, one at the beginning, people were wondering in general, why do you you What do you use for experimental data? Because you probably can't use any crystal data or this for disorder protein. Yeah. >> Yeah, so the we have basically only used solution experimental data, and that would mean in almost, you know, close to 100% either small angle X-ray scattering or NMR spectroscopy data. For the scattering data, if people all deposited the data, we could use the scattering curves. Often, we have to resort to using the radius of gyration. For NMR, we mostly use paramagnetic relaxation enhancement, although we've also done some work using chemical shifts. We've done some work, not that I showed you on optimizing against single molecule fluorescent experiment data. But basically, always we use solution data that reports on the conformational properties, at least in the stuff that I presented here here today. >> Oh, thank you. So, another question is What would you consider the main limitation of AlphaFold like approach when applied to intrinsically disordered protein? This is in the disordered regions, sorry. And there is a following up question in the contest still of the intrinsic disordered region, do you generally find apple or holo system to be the more informative for understanding their functional behavior? >> Yeah, so if you take the first question first, I I think that there is, you know, multiple layers to this. I I think that, you know, if we take AlphaFold-like models, it's a very broad term. I think the main limitation that, you know, in many of the applications is that they are, you know, they are trained with a loss function against a specific structure. And that can cause all sorts of problems if you have something that has a conformational ensemble. And so what I tried to explain is that, at least the way that we've done it is that we find it more relevant to target directly the experimental data. And of course, there are some approaches where you can optimize models such as AlphaFold, not by targeting XYZ coordinates, but by propagating all the way back to the, for example, the crystal data, but this is not so common anymore. And then of course, the models themselves need somehow to generate diversity. And you know, there are now lots of deep learning models that actually can generate diversity, but that has to be explicit in there also. So you need to have a model that generates diversity, and then your loss function somehow needs to, you know, include this diversity when comparing to whatever training data you have. As for the kinds of systems, you know, for most of the things we've looked at here, we really looked at at things that are very dynamic. So we've looked at the most disordered parts of it, and so we don't really, you know, we don't really focus so much on whether the folded domains are in an apo or holo state. And so for example, for everything I've showed you with folded domains, we basically freeze the folded domain in whatever structure we think is most relevant for the conditions under which the experimental data has been generated. So, if the small angle x-ray scattering data is generated for the apo protein, we simulate the apo protein. If it's generated with the holo protein, we will try to simulate the holo protein. But, we keep the folded domains more or less rigid. >> Okay. Thank you very much. We have Sorry. I just said Uh we we have another question. I thought we have a lot of questions, but I just pick up another one that is considering intrinsically disordered protein have more charged residue, how use of a model predominant on hydrophobic interaction can be justified? >> Yes. I mean, that's a a great point. So, in the first model that I showed you, there was no electrostatics, and that was clearly you know, suboptimal. In the model that we now have, the Calvados model, there is electrostatics. So, charged residues do exist. We treat them very simply. So, we don't have you know, you know, residues have a charge, and so we don't represent the charge distribution. If they are close to their pKa, we can give them a fractional charge. So, there's already a series of approximations that we put in there. And then these charge-charge interactions are are represented with sort of a screened Coulomb potential with a dielectric constant that we set to be a constant depending on the, you know, dielectric constant of solvent with a Debye-Hückel term. And so, we do have electrostatics in there, but it's not very refined. And I think there are many cases where this will fail. Um I will say I've been relatively surprised in how well it works and this is not just what we found, other people have found this also. So, I showed you a couple of of examples of both other people's data and our own data where we can take a sequence and we can move around the amino acids including moving around the charged amino acids to change the compaction and the model actually captures this very well. It doesn't capture particularly well salt dependency of compaction. It does to some extent. So, there are clearly some limitations in how we treat electrostatics and of course, if you have environments that have different effective dielectric constants, it will also fail. But, we do include it to, you know, to a level that I think is comparable in resolution to other parts of the model and I think we balance the electrostatics and the sort of hydrophobicity terms relatively reasonably in the model. >> Yes, thank you. I have a We have a There are one other question one question on the multi-domain model. Uh, do you do you put very strong harmonic restraint in your folded domain or does your model also allow to some fluctuation conformational change within the folded region? >> Um, the answer to to both questions is yes. You know, so in the Calpha 3 model, we have harmonic constraints in the AlphaFold Calpha 3 model, we have to use a more like, you know, Lennard-Jones like a 12-10 potential as, you know, is more commonly used in in in in, you know, what originally was called Go models or more general class of structure-based models. Um, I will say that we typically set them to be relatively strong um, to to sure that we keep the folded domains constant, we optimize this against the experimental data. You know, some years ago we published a paper where we combined go models with other physical force fields, sort of try to you know, model folding unfolding equilibria. And in principle, we could do this here also. Um, but in this case we said we we would rather err a little bit on the on the side of you know, keeping things restrained. So, we do model some dynamics, but I will say that the way that we model the dynamics is mostly that if we have loops, they are more flexible than secondary structure elements, and we don't try to optimize, you know, local fluctuations particularly accurately. If you are interested in some specific system, there is quite a bit of room of of of of treating these parameters, but for single parameter that works across all proteins, we went for something that is relatively conservative. >> Okay, thank you. And there is another question, um, that is uh, do you see any bias in the persistent length of conservative intrinsic disordered region versus not conservative linker pro linker versus protein that are full intrinsic disordered region that has full disint- intrinsic disordered regions, I think. >> I I think we've not really analyzed this in any way, so I will just say again, everything uh, that I showed is publicly available. All the thousands of simulations are available for anyone to analyze. I will say that I while I started by saying that 20 years ago we had a local potential that I think we optimized relatively well in the current model, we actually don't. And even that original local model was not sequence dependent. So, I'm not sure how well we actually model sequence variation in persistent length beyond what is governed by these hydrophobic interaction parameters? So, there is and other people have done this. There is room for improvement on developing models to capture this that can then potentially address the question that's being asked. So, it's a great question. We've not looked at it, but I would also not be sure whether the current model has the accuracy that could really address this. >> Yeah. Thank you very much. And how do you compare atomistic molecular dynamics with Calvados constrained molecular dynamics, I guess, for monomeric study of protein? >> They have different use cases. We still do some all-atom MD of disordered proteins. They are very, very difficult to converge. When you go beyond, let's say, 30, 40, 50 residues, you really need to spend quite a bit of computational resources and and, you know, use non-trivial sampling methods to converge them. But, of course, they give you huge amount of detail that our coarse-grained models and other coarse-grained models don't give us. And so, for example, local structural properties, coupling between local and structural properties, detailed interactions between side chains, um all of this you get from all-atom MD. I don't think that all-atom MD as it is now is actually more accurate than the coarse-grained models for measuring global structure. So, I think that if we could converge all-atom MD for these very long IDPs that we're simulating and we unfortunately can't, I don't think they would agree better with experiments, but of course, they agree much better with experiments on local properties and all um and or, you know, giving you the chemical detail that is often necessary for studying interactions with small molecules or or other proteins. There are of course intermediate level models. There's Martini-like models that have some reduced coarse-graining compared to, you know, it's more coarse-grained than all-atom, but less coarse-grained than Calvados. Or there is sort of all-atom implicit solvent models. And there's a bunch of models for IDPs. There's Absinthe, we developed one called EF1SB. And there are several models that can do these things that, you know, sit in this space between. They do different things, and I think it's a balance of, you know, how much computational resources you want to apply, how broad you want to go in sequence space, versus how much do you care about a specific system. >> Thank you. I reformulated the last question that I will do. It's What's the question was concerning nucleic acid, so DNA or RNA in the Calvados framework. >> Yes. So, we have a published model of what we'll call disordered RNA. And we call it disordered RNA because it has no secondary structure. Each, you know, nucleotide is represented by two beads, sort of a backbone bead and a as a base bead. There is no sequence, so it's really a flexible polymer that we optimize sort of the local potential to match all-atom MD. But of course, given that there's no sequence, it's relatively rough. And then the non-bonded terms are also optimized to give us global conformational properties. It's very simple, so there's a huge amount of things that it cannot do. But for some things, it actually does reasonably well. We then also have unpublished work that we hopefully will will be able to share soon, where we can do similar to what we've done for the Calvados 3 and alpha Calvados model, where we put on harmonic constraints to keep secondary structure either in DNA or in RNA um to study new systems where where where you know duplex RNA or DNA would be important. Again, these models are currently not sequence specific. They keep the structure rigid. We tune the parameters to get the persistent length and other interactions roughly correct, but they are very primitive and I would also say that they are you know much more primitive than other models out there um and they're more primitive than the protein part of Cavas, but they may be good enough to address some questions. So, if you want to address questions related to average stickiness or charges or some kind of heterotypic phase separation, maybe it's good enough. And you know, we'll we'll keep pushing these models, but I think it's also important to remember that there are limits to all models, be it all-atom MD or different levels of coarse-grained models, and there are just some things that this kind of coarse-grained will never be able to capture. And so, there's a limit to how much, you know, chemical detail that we can reasonably try to put into a model that's as simple as the models that I've been presenting. So, we we have models for nucleic acids. They are okay for some things and terrible for other things. And hopefully, we try to write in our papers what we think that they can do and what they cannot do. And if you're not pleased, then let us know. We'll try to do better. >> Thank you very much, Kerstin. Thank you. And there are still a lot of question. I invite all the attendees that ask question to use the Ask by Excel forum that I put in the presentation to ask the the question that you have here to to type your question there, so Kerstin can address those question. They will answer question in the next week in the forum. So, I thank Thank for attending. I thank you Kerstin. It a fantastic talk. You got a lot of nice talk mentioned from that attendees. They were all enthusiastic. So, thank you very much for join us for this BioExcel webinar. Thank you to for all attendees for being with us and for the question. I thank you Autumn Richard for join also this section and I will remind all of you that BioExcel has a BioExcel conference coming in September. Please register. The deadline end of May is the deadline for the early birth and presentation of abstract. So, please join it. You find all the details on BioExcel webpage. I can I can also share now the Maybe it's easy if I share my presentation so you can see the detail. Just give me a moment. Uh Sorry. So, that is uh So, that is the details. So, the barcode is there. So, you can join us. The 31st of May is the deadline for the early birth. And please join us in Brno in September this year. And then I wish all the other we stop here the spring section of the BioExcel webinar and we will have a new section coming up in autumn. So, thank you very much everybody. Bye-bye.