Submind YouTube summaries
Thumbnail for Bioexcel webinar #94: AQUA-DUCT - Solvent Tracking Software, From Protein Engineering to Drug Design

Bioexcel webinar #94: AQUA-DUCT - Solvent Tracking Software, From Protein Engineering to Drug Design

Watch on YouTube

Video summary

The webinar introduces AQUA-DUCT, a specialized solvent tracking software developed by the Tunneling Group at the Silesian University of Technology in Poland, designed to overcome the limitations of traditional geometry-based methods like Caver or Mo. Unlike these older tools that rely solely on static structures, AQUA-DUCT analyzes small molecule transport within proteins by actively tracking water molecules and other probes during molecular dynamics simulations. This dynamic approach allows the software to map transient tunnels and capture asymmetric features that static models miss, focusing exclusively on molecules inside a defined region of interest while ignoring external solvent. The system utilizes a color-coded visualization logic in PyMOL or its graphical interface, Kraken, to distinguish between entry points, exit routes, internal behaviors, and movements between cavities, all supported by statistical overviews of flow intensity and directionality. Beyond simple tracking, the software employs a grid-based approach combined with the Boltzmann equation to calculate transport barriers, identify gating residues, and locate regions of high solvent density known as hotspots where molecules may become trapped. These capabilities were demonstrated through diverse applications, including the analysis of epoxide hydrolases to reveal evolutionary conservation in tunnel networks and specific gating residues like isoleucines and phenylalanine. The tool also elucidated substrate specificity in formate dehydrogenase by showing how mutations creating salt bridges alter flow dynamics, and it aids drug design by using cosolvents as probes to map hydrophobic and hydrophilic properties within protein interiors, helping researchers design inhibitors that fit specific cavities to address clinical trial failures. For effective operation, AQUA-DUCT requires molecular dynamics trajectories with fixed periodic boundary conditions centered on the protein, proper hydration, and a sampling ratio of at least 1 ps per nanosecond, though users must carefully define the scope of interest as larger areas yield longer calculated paths. While the software is powerful, its computational cost scales with the number of tracked molecules, meaning large systems or long trajectories may require trajectory splitting to remain feasible, and tracking solutes like glycerol or glucose statistically significant results generally requires a system containing more than approximately 20 molecules unless compensated by extensive simulations. The developers are currently working on automating pharmacophore feature selection for drug design and note that while the tool supports membrane proteins, capturing rare events in hydrophilic environments still demands sufficient sampling. The webinar concluded by addressing practical limitations and future directions, noting that questions regarding specific solute tracking scenarios should be directed to the BioExcel forum for answers within the coming week. The presentation also highlighted upcoming educational opportunities, including a session on DynaPym scheduled for the end of the current month and another focused on multiscale simulation of biomembranes set for April 14th. Overall, AQUA-DUCT represents a significant advancement in understanding protein transport mechanisms by integrating chemical properties and dynamic behavior into a robust analytical framework that bridges the gap between protein engineering and drug design.
Read the full video transcript
Welcome to the BioExcel webinar Aquaduct solvent tracking software from protein engineering to drug design. Uh I'm Otto Anderson and usually this is hosted by Alessandra Villa, but she was unavailable today, so I'll be doing my best. And I have uh my colleague Lakshmana with me today also here. Hello everyone. Okay, this webinar will be recorded and uh it will be made available afterwards on the BioExcel website. And for questions, we will use the Q&A function. So uh during the presentations, you can ask your questions with the Q&A button found in Zoom and after Yeah, after the presentations, we will uh pick uh the questions from there and also allow you to ask them yourself if you have a microphone. It makes it easy to interact. Um and if there's still questions that are unanswered afterwards, you can ask them in uh the BioExcel forum. And today's presenters uh Arthur Gora. Uh Arthur is a professor at the Silesian University of Technology in uh Gliwice, Poland, and head of the tunneling group where Aquaduct, a software for small molecules tracking, was developed. And then we have uh Weronika Bajgrowicz-Czakon. Weronika is a PhD candidate at the Silesian University of Technology and a member of the tunneling group. Uh she's been involved in testing the Aquaduct software and is currently focusing on exploring its potential applications in drug design projects. And also uh Maria Błazowska. Uh Maria obtained her PhD from the Silesian University of Technology, where she studied uh, molecular aspects of protein regulation with a focus on the role of water molecules as mediators of inter-molecular interactions. Uh, currently she's a post-doctoral researcher in the group of Francesco Luigi Gervasio of in the University of Geneva. And that's it for me and let's uh, get to the presentations itself. Thank you, uh, Otto for introduction. I guess you now see my uh, full uh, screen. Yes. Excellent. So, one more time, thank you for introduction. Um, uh, I would like to start from acknowledgement of the uh, members who contributed to development of Aqueduct and also other groups of my members, funding and so on. This would be difficult to process. On the video you have seen also gating phenomena where I was stopped by gating residues. This is part of our job what we are doing. Here, uh, the presentation will be um, divided into three parts. I will present um, introduction based basic idea and applicability of our software. Then, uh, Maria will present how to run and um, the examples of the analysis with the use of Aqueduct will be presented by Veronica. Uh, so just um, if you imagine the protein from point of view of the small molecules of the substrate or the inhibitor or any other molecules, you can think that the the molecule see the uh, surface which is surrounding the the protein, but also it could be just other um, um around. So, it can see the surface or the interior of the protein, yeah? Both of them they are important, of course, from our aspect. The question is how to how to look into the second one. How to look into the interior properties? And there was a lot of um uh tools developed uh before based on geometry um based approach, uh such as Caver Mo or others, and they work very nice, and they are very uh good for many of things. They can detect the tunnels, they can detect cavities. However, they have some some some limitation. One of the limitation is that since they are using uh usually um spherical um spheres to approximate the tunnels or cavities, then uh you have problems with facing uh and asymmetric features both in case of the interior cavities as um tunnel entrance. The second thing is that since they are based on the mathematical algorithms, they don't recognize the chemistry. So, they the same um the tunnel with the same diameter uh which will has the charge or not charge residues are treated in exactly in the same way. The other thing is that even if you can expand them to analyze molecular dynamic simulation, then in a case of transient tunnels, they won't be detectable because you need to see the pathways uh from the beginning till end. Uh so, still from, let's say, active site to the surface of protein within one frame. If there is some disconnection uh in a in a snapshot of the MD simulation, then the tunnel is not detected. So, we were thinking how to um how to solve this problem, and one of the uh our understanding was that if we look into MD simulation, and this is the frozen image from one snapshot, then you can see that water molecules pretty well penetrate also protein core. So, then you can find them in the uh inside in active site in in in places between so in the in the in the tunnel. Um so, we had an idea to use this possibility to search for the tunnels and describe them. However, when we look into single water molecule behavior, this is 50 nanoseconds of simulation and just single water molecule, you see that the trajectory of this molecule is pretty pretty complex. So, in one moment we were facing the problem how to solve it, how to approximate if we want really to use this method since we have more than sometimes more than dozen thousand of of small molecules in our system. And how this is how we we develop the Aqueducts, so the software to track and to solve this this problem. The basic the concept of the tool was very simplified. So, um Sorry. So, if you So, we are not considering water molecules which are out of the protein surface. Yeah, if they are in the surrounding that we neglect them completely. Then we minimize the search of the water molecules only for those one which are in the object. Object is something what we call the place of our interest. So, for example, let's say that this is active site cavity. Then you can see that we use some color coding. So, the red is And then we are tracking only those molecules water molecules or some small molecules which were detected in in object. So, in this case in active site. And we are tracking them how they are leaving and how they get into the object or the active site. We use the color coding so you see that there is red pathways show entering molecules, blue the leaving molecules, the green one depict how the molecules behave within the active site or the object of our interest. It can be also co-factors site or whatever else and the yellow is the other possibility and that's the molecules which are penetrating interior, leaving object and coming back to the object but they are not leaving the protein. So it means they they are just going to other internal cavities and and and and and and provide some description of the of the interior. The two other things which are important it's are inlets. So inlets are the points locations in a in a space which are on the border of the convex hole which describe the protein surface and they are just informing us where the entry or the exit of the tunnel is located. How this concept look into in practice? That's the example of single water molecule pathway traveling through cytochrome P450 where you can see that in one location was entered then it was pretty long time staying in one location because of the gating residue which was closing entry of this water molecule further. Once the molecule the the residue side chain moved then water molecule moved forward it reached the active site a heme and then leave by other tunnel located in other place of the of the protein. Um this can you analyze for all water molecules and then you have statistical overview of all information of all water molecules which were traveling through your protein, through your active site during the entire simulation. You can visualize this on the left side you will see the shape of the tunnels and you you see it's not the circular, it's really representing the dynamic of the um of the protein and also the colors in the identify the intensity of the number of the inlets. So then you you see where majority is um is is is coming from or leaving. And on the right side what you see, you see again it's a cytochrome, you see the picture of um five, uh let's say um tunnels, uh which I mean five entry points to the protein um and and then you you see the statistical interview how many water molecules were detecting traveling through particular entry. And the um the picture below the the circular picture below, it's maybe more complicated but it's very very informative because it's providing information about direction of the of the uh of the water traveling through the protein. So then you can say for example that from 2B uh tunnel through the entry of 2B, uh to tunnel number four, the 30 um the particular number 22% of water molecules were were traveling. Yeah, so you you you see the whole directions and you can see the uh the network of the tunnels, how they it contribute to the travels of the of the molecules. That's a dynamic picture, and you can also analyze the other features. So, you can also analyze the local distribution of um of uh water molecules. For that, we are using the the grid. So, then we are calculating number of water molecules uh or any other molecules uh during the entire simulation um in a cubics. Uh these cubics um then provide statistics, and then you can uh do several things. The one thing is that you can allocate them uh with the pathway, which usually for that you uh simplify. And then, based on the statistical distribution and using Boltzmann equation, you can calculate approximate the barrier of uh um the the barriers of the the the the location of the um most difficult um places to to move uh water to transport water molecule, then you can identify uh gating residues or other obstacles which which um contribute to the um uh control of the of the transport of the water molecules. And the second thing is that uh you can also identify uh the location which are the most uh most often occupied by water molecules. So, the density the local density of the solvent is uh is uh the highest. And that's we are calling the hotspots, and they of course uh represent two uh scenario. The one is where uh molecules were trapped, and the other where they strongly bind. And um that's uh something what we can use in a variety application. So, uh if you ask how we can use uh this this type of analysis, uh then the first is what I show already is transport analysis. So then you can depict the information about the transportation of small molecules through your protein. In many aspects, you can identify also the very rare events. So the pathways like here the ones where only few molecules go through and that can be used for engineering. Then you can use this information for structural analysis of the of the proteins. This will be one of the example of the applicability and the show by Veronica. You can a little bit understand evolution of the protein especially in if if the tunnels are important due to the the protein uh um catalytic machinery and you can also use it for drug design or inhibitors design. Here specifically you can use not only water molecules, but you can use also a simulation in cosolvents and then you can use every small molecules as a specific molecular probe testing different properties. Then you can build pharmacophores and based on on on hotspots and try to design inhibitors which match the pharmacophore. And you can use also for protein engineering and testing why particular mutations cause difference in activity especially where it it it reflects some issue with transportation of the substrate product release or also with accessibility of the water. What is more what is important also is that a product is system independent. It's ligand independent and it's highly flexible. There is several models how you can use it to merge different simulation together or to or to divide single simulation into counterparts just to see how the dynamics is it's it's modifying the transportation phenomena within protein. And then then the the three examples which will be shown later on comes from our study one about epoxide hydrolases when we tested aqueduct on seven different epoxide hydrolases and as you can see here on this photo on the on the slide they were very diversified in case of the how how network and how accessibility of the access site it's present. And you can see that the mammalian has very rich number of tunnels entries where the the plants are preferential one tunnel which is in main domain of epoxide hydrolase and instead bacteria has the entry the main entry and the the primary entry through cup domain. So it's kind of evolutionary conserved feature. And this one of this epoxide hydrolase from Solanum tuberosum will be used as example how to analyze the results to to describe structure of of this of this enzyme and it will be shown by Veronica. The second example shown by Veronica will be the work recent work of Agata Raczyńska who unfortunately couldn't join us today. And it was about understanding of substrate specificity of the pathways for transportation of carbon dioxide and uh formate through uh tunnels in times independent formate dehydrogenase. And also how does it work with the mutants and why this mutant is specific. And the last one will be example about drug design and here we again will go back to epoxide hydrolase and this is our study where we use several cosolvents to map hotspots in the interior and try to to find the new cavities, new properties or new residues inside protein interior which could be targeted. The reason of this study was that most I mean all of the inhibitors uh uh so far developed they didn't pass the clinical trials so we were searching for some inhibitors with unique features and and it's will be also one of the example on the last last example. And that's I think all from my side. I would like only to share that uh um most of the information will find of our Aqueduct website dedicated for the software and you can also find um in the Living Journal of Computational Molecular Science the the tutorial which is very very complex but it's going through all um all steps of analysis and also of preparation for the calculation which we now show in shorter version and it will be shown by Maria Błufka. Yes, so let's move to how to use Aqueduct. Next slide. And as you can see here this is like the global picture how the Aqueduct software looks like. What we need for the Aqueduct to run the calculations is first of all the trajectory from our molecular dynamics simulation together with the topology file and also some config file where we put all the necessary information. What is important here that Aqueduct use two drivers. The first one called valve which is the main driver to perform the whole analysis of the trajectory the selected ligand in the simulations and then the complementary pond driver which does this advanced analysis based on the local distribution and local density analysis that are to mentioned and this driver pond uses all the information that were previously computed by the valve. So it means that for all the distribution analysis that you want to perform to obtain the information about the density, the pockets, the hotspots, first you need to track your molecules of interest. We also have a small application that helps to gather and visualize the results and what also Veronica show later is that the ultimate output can be visualized in PyMol as a PyMol session. So let's move further. What is really important and what we want to stress out here is like what are the requirements that you need to follow before running the Aqueduct calculation. So first of all what you need is the trajectory as I said that contain information about both the macromolecule of the interest here usually the proteins and the small molecules that are present there. So usually water molecules but as we can say it can be ligand independent so we can track ion, substrate, cosolvents, whatever you want and whatever is present in your simulation files. And then before running the uh calculations, you should really fix your trajectory files by fixing the periodic boundary conditions. The protein the macromolecule of interest should really stay in the center of the box and all the movements should be reduced. Then uh what is really important, especially for the analysis of water molecules, it is how frequent we save our um movements to capture them the movements properly. So, we know that right now we have longer and longer trajectories, but we should really point out that the user should aim for the 1 picosecond per 1 nanosecond ratio and not exceed the 5 picosecond of sampling because otherwise it would be really hard to capture um the relevant movements of the molecules. And again, here because like the concept of the Aqueduct was strongly connected with analysis of the water molecules, please be aware of the proper protein hydration before even running the MD simulation. There are plenty of tools that are available to place water molecules properly inside the protein cavities or run sufficient equilibration of your system to enable to properly penetrate the solvent molecules the the protein interior. And what is also important is that for running the Aqueduct calculation, we require a functional um Python environment, preferably Python 3.9, and we strongly suggest to do all the Aqueduct calculation without the within the virtual environment. Installation of the Aqueduct is pretty straightforward. It's described also on our website. So, uh it's really easy to install the AQUEDUCT and then follow all the tutorials and run your calculations. Okay, we can go further. And here what is important and what also Arthur mentioned is like the definition of our macromolecule and the object of the interest. So, object is like the part which interesting for us the most because within this area that we define, we track the the molecules. Uh and usually it is uh connected with some active site or cavity within the protein molecules. So, the definition of that uh should be should be done with uh careful consideration and you should really also define like which molecules you want to track that are passing this object of the interest. And as Arthur said, we are restricting our um investigation area to the scope, so the area within the analyzed system within the molecules should be traced. We can go further. And for you just to mention is that the AQUEDUCT can be run in two different ways. For guys, you can use either the graphical user interface which should um facilitate the running the calculation and also preparing the configuration file which is necessary to run the calculation, but you can also do it just in the terminal and then just by um manually editing the configuration file. For the new users, we strongly recommend going through the graphical user interface because then just uh it guide you how to properly um prepare your configuration file. We have different levels of preparation of the configuration file starting from like very easy mode when we just trace uh molecules of interest going through the more advanced modes where you can also perform more advanced analysis such as hotspot analysis or pocket analysis. And then we have also this small graphical user interface called Kraken for the visualization of our results where you just upload the data that you obtained from the aquaduct and you can obtain the um uh nice visualization of your results. And here what is really important and what we want to stress out uh is the proper definition of the scope because by the defining or redefining the scope uh we can obtain like different information about the trimming of the paths and paths are important because of the tracing of the molecule and like area where the molecules entered or uh leave the macromolecule of interest. And the general rule here is that the paths are really limited to the scope area. So the bigger the scope is, the longer are the paths. So you can imagine that if you analyze the protein, so if you want to um have your inlet so like the the positions where the uh, molecules enter or leave the uh, scope to be located at the, uh, greater distance, then you should define the, scope as the entire, uh, protein. And if you want to go narrower and narrower, you can, uh, go either through the backbone definition or even the main C alpha, uh, definition, uh, to obtain shorter paths. And here is the example on the, uh, left side you can see, uh, how the, um, how the, uh, scope is defined as a, as a protein. So, those, uh, purple lines shows the definition and those small X's show, uh, where the, where the inlets, uh, will be located. While with this, um, uh, more deeper definition, uh, when the scope is defined as a, uh, backbone, the, the inlets, uh, will be, um, will be displayed, uh, deeper, uh, in a protein. We have also additional, uh, procedure for trimming the pathways, uh, which was developed, uh, together with the, um, uh, Aqueduct. It's called, uh, AutoBarber procedure. And, uh, it, uh, works by trimming some, uh, pathways accordingly. Then, uh, as Arthur said, we can have this, uh, information where the, the molecules point out the entrance or the exit area uh, to the interior of the protein. And, basically, this is this golden, uh, spheres area. We have many such of the points. And usually, we want to de-cluster them accordingly. Usually based, on the, uh, specific regions, uh, where those molecules, uh, enter. And there are, uh, many ways to divide one big cluster into, uh, smaller pieces. In Aqueduct, we implemented different methods for that. And if you go further, Arthur, you can see that based on um a choice that user is that that the user is doing, we can obtain like slightly different pictures where the clusters will be located. So, if we combine that with the information from the paths, we can really specifically cluster the inlets accordingly to the entrance and exit area. And if we go further, we have this like quantitative analysis that we can say and Arthur already mentioned a bit about that. So, this is just like the graphical visualization of the of the results about the cluster size, the relative flow between the between the cluster, the intramolecular flow, the directionality of the of the flow, and also the time within the simulation where particular molecules entered our scope and objects. And if we go further, we also have this information from the distribution analysis when we can analyze the area that were visited by our tracked molecules. We can obtain information about the the volumes, the space which was penetrated very often. This is the picture A where the molecule spent sufficient enough of time. But, we also can obtain the information about the maximal volume that was penetrated by the tracked water or other molecules. And based on the density information, as also Arthur said, we can obtain this information about the hot spots. So, the points within the macromolecules with the highest local density. Uh so, basically we can obtain the information about specific location where the molecules of interest were attracted, trapped. We can see that they differ um in size here. So, the biggest molecules this spheres means that the highest density was exactly in this specific area. And last but not least, as also Arthur said, we can perform this analysis for the um not only one type of molecule present in our system. If we have the simulation with the cosolvents, and here we define cosolvent as any um um low molecular mass compound that can be mixed with water molecules, and they act as those like very specific probes. And here we can clearly distinguish like how different they can penetrate the interior of the protein. For instance, different parts than water molecules. And it can give us like the insight uh that we could put particular focus on very specific parts of the protein interior um based on the type of the cosolvent or solvent that we were tracking. And with that we can go to the last slide where we just show that how you can visualize the results. It's very simple but by typing the command to open the script for the visualization and here in the in the PyMOL you can see the loading of the uh all the results when we have information about the protein, the clusters, and also all the paths between uh the clusters and also like uh single um paths of every single um water molecule that was tracked. Here. And with that we will go to Veronica and practical examples. Yes. >> Yes. So, I hope you can see my screen. Uh as Arthur mentioned, we will go through three different example of usage of Aqueduct and we will focus on the results visualized in PyMOL. Uh for today we selected for you three examples. So, simple water tracking, then substrate tracking, and mutation effect at and the end uh drug design example. So, we will start from water tracking. For this we use uh epoxide hydrolase from potato. In this enzyme water is very important for the reaction of the hydrolysis. So, Maria introduced you to the definition all of the things we will see. So, I will not say that once again, but uh here you can see uh the scope. Scope is defined here as a protein. And you can see object. Object is defined as a sphere uh of the amino acid which create the active site. So, how you will visualize the results, how you will look into your results, it's up to you how you prefer to look on the protein. Uh we prefer to look uh on the subs- surface of the protein. So, if you will visualize the surface and you will open the cluster, you can see where is your tunnels. So, here for example, you can see one tunnel and you can see that it's a hole. So, you are sure that here should be some tunnel, but if you will look, some tunnels can be deeper a bit inside and they are not visible on the surface because some tunnels can be closed in the selected frame. So, you can use the transparency mode on the surface and multi-layer is now on. So, you will see the interior cavities and your surface in one time and this is how we prefer to look into the session. So, here you can see that in this epoxide hydrolase, we have one main cavity and we have the some tunnels here. The third one is here and some inlets here at from the other side. So, based on analysis, you can see how big your cluster are and you can see and say that, "Okay, this tunnel is the main tunnel because the flow is big here and the rest are used more less frequently." You can also look into the paths. Here you can see single row paths for single water molecule. So, as Arthur said, you can see that sometimes the water is trapped in some parts. So, you can analyze all of the places. You can identify the amino acid which are trapping the water, but you can also look uh on the path and see where water enter to the protein when leave protein. So, for example, here you have the red color. So, it means that the water enter by this small two inlets cluster and it leaves the protein through this one bigger cluster. And if you will look into the residues, we are close here, you will see that here we have as asparagine, uh some glutamic acid and leucine. So, they are responsible for this opening of the small hole. And you can also analyze here, for example, the bigger tunnel. So, you will see that here we have two isoleucines which can work as a gate here, but also we have two, uh, aromatic residues, phenylalanine and tyrosine, which are flexible and they can, uh, they move and they are responsible for opening of this second smaller tunnel. It's very worth to look into your MD data to check which amino acid are flexible and connect the results of Aqueduct with the flexibility because maybe you will find some important gating residues and, uh, for example, residues which you will, uh, mutate in the future steps if you will use, uh, the results for enzyme engineering, for example. But, uh, besides such a single path look, look, you can look into the results in the more more global way. So, here you can visualize not only a single paths, but the whole paths and, for example, you can see, uh, the paths between the tunnels. So, for example, here you can see, uh, two four, it means that the water enter by the cluster number two and the leave protein cluster number four. You can say see the opposite side, so for cluster number four to two. And if you will open all of them, you can see that you have a lot of paths inside and some are more dense, some less. So, that's why how we look, uh, for searching the density and hotspots. But, when we move to the hotspots, here are some uh, important things. You can also visualize paths of the molecules which enter to the protein and stay inside the end of simulation of the paths with for the molecules which were inside from the beginning of the simulation and escape through different tunnels. Of course, you can see the hotspot, so the most dense part in the protein. And here, for example, you can look into this red hotspot, and you can see that it's close to the catalytic histidine. So, it means that somewhere here it's very important water, very often, which is required for the reaction of the hydrolysis. And you can visualize pockets. So, we are using mesh visualization, and here you can see the inner pocket, so the water was most during the simulation was likely in this cavity, and you can see outer, so the whole space for the water, and you can see the differences between the pockets, so the inner is smaller than outer, and this results can show you very a lot of very useful information because you can compare it with the cavity for your crystal structure, for example, and then you will see that the cavity during the Aqueduct analysis is much bigger, can try to find new spaces in your protein, new cavities, and maybe try to put new inhibitors in the new space which you defined. Each Aqueduct results give you a two files which are not visualized in PyMOL, so txt file and csv file. And within this file, you can find a lot of very useful statistic data. We will go through this text file, so you can see first you can see all of your config file configurations. So, you can find how you define the object scope and so on. So, it's a lot of such a data. Then, you can see how much frame you used in your simulation, which type of molecules were used to trace. And here you can see how many water residues are found. So, it's 936. And for this number of residue, we have 1,299 two paths. And the number of inlets is 2,532. And you can think why we have more inlets than residues to trace. So, some residues can enter the protein, then go out, go again inside the protein, and so on. That's why we have much more inlets than residues to trace. Of course, you can find all of the statistic about clusters. So, the clusterization procedure, the cluster data, so how many inlets you can find in your cluster, how many incoming, outgoing, and different type of statistic in the tables. They are all described. Here is one big table. So, here in this table, you can see a lot very useful statistic for single water molecule based on ID of this molecule. You can even found the time, so in the frame when the molecule entered the active site or entered the scope. And when molecule leaves from the active site. So, we will go to the second example. Here we have formate dehydrogenase is the enzyme that perform the reaction of reduce CO2 to formate and within this example we trace the substrate so we trace CO2 we also trace formate and we prepare some mutation. So here you can see as a mesh visualize the tunnels detected by Caver. You can see that Caver detect two very narrow tunnels and if you will look into the paths you can see that for example CO2 it's using two tunnels. Here the smaller tunnel and the big tunnel which is here and what is important here you can still see here the Caver tunnel is very narrow but CO2 use much bigger cavity for the transport to the active site. And if we will just look into single paths you can see that some example of enter of CO2 to the active site and escape from the one big cluster but if we will go to different paths for example here you can see that this tunnel probably were closed in that time. So CO2 molecule decide to just go back the same tunnel and escape and what is very curious here that this tunnel is very narrow. Here you can even see the space between the cavities. So if you will look into amino acid in this area you will see that here we have a gating residue is tryptophan and is one arginine residue and we decided to mutate this residues. So if you will look into the mutation here is glutamic acid and you will see that glutamic acid can create of arginine the salt bridge. So if you will look how CO2 is going through the protein within this mutation, so we will see that whole transport is almost by this one cluster. The second cluster is used into the smaller way. So, it's just two inlets here. And uh here we can also look into the single path and we will see that here, close to this salt bridge, water was trapped CO2 was trapped for some part of simulation and then escape when this bridge disappears. So, even such a change like small bridge, which is not so strong interaction, can change the flow of the substrate or the solvent within the protein. And if we will go to the uh formate, we can see that uh here we have also these two tunnels detected by Caver. We will see that formate is using just the one main tunnel to go uh to the active site of the protein probably because of uh the size of the formate. Formate is bigger than CO2 and this second tunnel is just too narrow to be used uh for enter or leaving the molecule. And uh the last example it's uh example of human soluble epoxide hydrolase. Here you can see the structure of human soluble epoxide hydrolase with all of the crystallized uh inhibitors. So, this enzyme performed the hydrolysis reaction of very important epoxyeicosatrienoic acid, which can be used as uh potential non-inflammatory compounds in the treatment of many diseases. And for today, we don't have any good drug on the market, which is working actually no drug pass uh clinical trials. So, here we used Aqueduct with cosolvents. So, you can see that we have different type of cosolvents and we can describe the pharmacophore features to each of the molecules. So, we have acetonitrile, which is the hydrogen bond donor and hydrophobic compound. We have dimethyl sulfoxide, which is hydrogen bond acceptor. It's here. Green one, we have methanol. Methanol is hydrogen bond donor. Phenol, which is also aromatic. And here it's not so good visible. Phenol, it's here. So, small hotspots. Then we have urea, which is the bond donor and acceptor. And of course, water, which is also hydrogen bond donor and acceptor. And what is curious here that if you will look into detail into the inhibitors, you can see that if you have a group which are able to create hydrogen bonds, so they have nitrogen or oxygen, water hotspots are very close. For example, here we have more hydrophobic part and here we can see the hydrophobic part of inhibitors. Here where is this phenol, you can see a lot of aromatic strings within the inhibitors. So, it's kind of the proof of our concept that we can see the description of the places based on the possible interaction within the protein. And now if we will switch off the inhibitors, we can see that we have a very nice map of potential interactions or even the description of the places. So, we can say a lot which for example part of the protein of the cavities more hydrophobic, hydrophilic, which prefer to have some aromatics part of the inhibitors. So, based on that, we can design the inhibitors or try to even find by virtual screening the compounds which will fit to our cavity based on the interaction which we can describe here. And now we are working to go more into the detail within this. So, we are trying to extend Aqueduct a little bit to describe better the properties of the cosolvents and to automatize the selection of the features for some drug design project. So, I think that's all from my side and I think that we still have a time for some questions. So, if there are Yes. >> one. Yeah. Yes, there was one. I lost the window now. Uh sorry, there was one regarding membrane proteins. So, I think Maria, you can elaborate about that. The question was if you can use Aqueduct for membrane protein and if there are any limitation for that. Um I mean, yeah, basically if you have such a big system like a membrane protein, um I suppose like I don't like GPCRs in certain inserted in some membrane like a transmembrane protein like a receptors. Obviously, you can use Aqueduct for this analysis. Like the way here is like the definition of again like the scope and the object of the of the interest. But and I think that that was something about the ion channels that uh you can also use that for the for the ions channels. Here like there is no problem with the definition itself. But the issue here is that uh you should have like uh If you have a very hydro uh hydrophilic environment and when it's really hard for instance for water molecules to uh pass, then you will have like very rare sampling in your MD simulation and then maybe it can be hard to um to describe like the particular events uh present in your system. But for instance for the ion channels you can track the um ions in in the channel and here it can be defined as a molecule to track and you can obtain all the information about the what the specific kind of ion is doing within your channel. Okay, there's more questions here. Thank you very much for the lecture and the cool results. I'm curious how long does the simulation and the analysis itself take? The simulation by itself it's just MD simulation so and the the aqueduct depends I mean calculation of aqueduct depends on how many water molecules in fact you tracked. So of course if you have a larger number then it's getting longer and longer and this is something what is difficult to um to uh to to paralyze. So, so we have some kind of problems if the system is too large or too long. Uh it it it can be a problem. Uh we have some kind of benchmark uh in publication which was done for 50 or 100 nanoseconds if I remember and number of different water molecules. Um what we truly recommend when you are running analysis is to pick um uh some smaller uh time uh or part of the simulation first just to really design properly the scope and the object uh just to test what settings are the best to answer your questions. And then truly to think twice um how to minimize the effort for calculation because the number of water I mean, once you are reaching really high number of of of molecules to track, then it starts to be long-term project. The largest system which we used was Maria in fact using it was TLR um part of the TLR receptor together with other protein and then we define the object in on the interface between these two proteins to to track the water molecules which participate in reaction of the of the loop cleavage. So, you can use very high system, but if you properly define the object and the scope, then you can minimize the cost of the calculation. You need to be really careful and smart. Exactly. Here, like you don't need to define the scope as an entire complex. You can focus only on the specific part and that would be something that we recommend to overcome the limitation with the number of atoms in your system. And sometimes it's better to use shorter trajectories in replicates than one very very long one. for our calculations. Okay. Then next one. Oh, you want to continue? Yeah, what are the limitations? So, the limitation is it's it's really the number of I mean, we didn't test what is the maximum number because it's really system dependent also the resources dependent. So, it's it's hard for me to answer this question. Basically, what you can do always you can treat system in a little bit different way and just divide simulation into small party small parts smaller parts if this very very large number of track molecules and then you can solve the issue. Regarding the updates, one thing which we are thinking and we are working on it's it's it's it's it's for pharmacal for design but this is kind of complicated problem and uh and and we are trying to solve it already 2 years. So, once we will solve for sure we will share to to community. Then there was a question regarding ramp and ligand exits and compare of uh Yeah, so this is this is possible. I mean, the power of Aqueduct is in statistic. Yeah, so if you have just a simple single molecule which you are tracking, then Aqueduct is not the best tool for that. Yeah, but if you want to to how water is assisting your substrate, then of course you can do comparative analysis, so you can run a product for one and the second event and try to to compare this data. We also built some feature which was tested on on one of the example exactly trying to calculate the barrier of the entry of the substrate, so then we were what we were doing we were in fact merging fragments from simulation from different simulations, so let's say that you have 10 simulation when some event occurs very rarely, but you pick only the the parts of the simulation when the event took place and emerge them together and analyze together. So this is also somehow possible one of the component, however, it is tricky and I think the most complicated job for aquaduct to to do. But for sure if you have long enough sampling for both states, so ligand bound and ligand dissociated, then you can do comparative study and for sure you will see difference in opening or closing if they are yeah or some conformational changes which will differentiate the the for example interior space. uh 300,000 We have time for one more question, so maybe you can read one aloud and then someone quickly and then I will have to close. um So regarding the last question, if there is any problem with solutes like glycerol, glucose or another another class um again it depends on how many uh, molecules you have. There is no problem for Aqueduct regarding what it is to track. Just to to be to have consistent or good results, you need to have simply more than let's say 20 20 molecules in a system. Yeah, if you have less than 20, then statistic will be very very small. Or you need to have very long simulation multiple replicas. All right. Any other questions, you can please put them in the BioExcel forum and the presenters will try to answer them the coming week. Yes, we will. Yeah, thank you for opportunity. >> I'll just take my this opportunity to tell about the upcoming webinars. There will be one in still at the end of this month about DynaPyn and then the next one is actually in April 14th, multiscale simulation of biomembranes. So, thank you to everyone. Thank you. >> Thank you.