Bioexcel webinar #94: AQUA-DUCT - Solvent Tracking Software, From Protein Engineering to Drug Design
Watch on YouTubeVideo summary
The webinar introduces AQUA-DUCT, a specialized solvent tracking software developed by the Tunneling Group at the Silesian University of Technology in Poland, designed to overcome the limitations of traditional geometry-based methods like Caver or Mo. Unlike these older tools that rely solely on static structures, AQUA-DUCT analyzes small molecule transport within proteins by actively tracking water molecules and other probes during molecular dynamics simulations. This dynamic approach allows the software to map transient tunnels and capture asymmetric features that static models miss, focusing exclusively on molecules inside a defined region of interest while ignoring external solvent. The system utilizes a color-coded visualization logic in PyMOL or its graphical interface, Kraken, to distinguish between entry points, exit routes, internal behaviors, and movements between cavities, all supported by statistical overviews of flow intensity and directionality.
Beyond simple tracking, the software employs a grid-based approach combined with the Boltzmann equation to calculate transport barriers, identify gating residues, and locate regions of high solvent density known as hotspots where molecules may become trapped. These capabilities were demonstrated through diverse applications, including the analysis of epoxide hydrolases to reveal evolutionary conservation in tunnel networks and specific gating residues like isoleucines and phenylalanine. The tool also elucidated substrate specificity in formate dehydrogenase by showing how mutations creating salt bridges alter flow dynamics, and it aids drug design by using cosolvents as probes to map hydrophobic and hydrophilic properties within protein interiors, helping researchers design inhibitors that fit specific cavities to address clinical trial failures.
For effective operation, AQUA-DUCT requires molecular dynamics trajectories with fixed periodic boundary conditions centered on the protein, proper hydration, and a sampling ratio of at least 1 ps per nanosecond, though users must carefully define the scope of interest as larger areas yield longer calculated paths. While the software is powerful, its computational cost scales with the number of tracked molecules, meaning large systems or long trajectories may require trajectory splitting to remain feasible, and tracking solutes like glycerol or glucose statistically significant results generally requires a system containing more than approximately 20 molecules unless compensated by extensive simulations. The developers are currently working on automating pharmacophore feature selection for drug design and note that while the tool supports membrane proteins, capturing rare events in hydrophilic environments still demands sufficient sampling.
The webinar concluded by addressing practical limitations and future directions, noting that questions regarding specific solute tracking scenarios should be directed to the BioExcel forum for answers within the coming week. The presentation also highlighted upcoming educational opportunities, including a session on DynaPym scheduled for the end of the current month and another focused on multiscale simulation of biomembranes set for April 14th. Overall, AQUA-DUCT represents a significant advancement in understanding protein transport mechanisms by integrating chemical properties and dynamic behavior into a robust analytical framework that bridges the gap between protein engineering and drug design.
Read the full video transcript
Welcome to the BioExcel webinar
Aquaduct solvent tracking software from
protein engineering to drug design.
Uh I'm Otto Anderson and usually this is
hosted by Alessandra Villa, but she was
unavailable today, so I'll be doing my
best.
And I have uh my colleague Lakshmana
with me today also here.
Hello everyone.
Okay, this webinar will be recorded and
uh it will be made available afterwards
on the BioExcel website.
And for questions, we will use the Q&A
function.
So uh during the presentations, you can
ask your questions with the Q&A
button found in Zoom and after
Yeah, after the presentations, we will
uh pick uh the questions from there and
also allow you to ask them yourself if
you have a microphone. It makes it easy
to interact.
Um and if there's still questions that
are unanswered afterwards, you can ask
them in uh the BioExcel forum.
And today's presenters
uh Arthur Gora. Uh Arthur is a professor
at the Silesian University of Technology
in uh Gliwice,
Poland, and head of the tunneling group
where Aquaduct, a software for small
molecules tracking, was developed.
And then we have uh Weronika
Bajgrowicz-Czakon.
Weronika is a PhD candidate at the
Silesian University of Technology and a
member of the tunneling group. Uh she's
been involved in testing the Aquaduct
software and is currently focusing on
exploring its potential applications in
drug design projects.
And also uh Maria Błazowska. Uh Maria
obtained her PhD from the Silesian
University of Technology,
where she studied uh, molecular aspects
of protein regulation with a focus on
the role of water molecules as mediators
of inter-molecular
interactions.
Uh, currently she's a post-doctoral
researcher in the group of Francesco
Luigi Gervasio of in the University of
Geneva.
And
that's it for me and let's uh,
get to the presentations itself.
Thank you, uh, Otto for introduction.
I guess you now see my uh, full uh,
screen.
Yes.
Excellent.
So, one more time, thank you for
introduction. Um,
uh,
I would like to start from
acknowledgement of the uh, members who
contributed to development of Aqueduct
and also other groups of my members,
funding and so on. This would be
difficult to process. On the video you
have seen also gating phenomena where I
was stopped by gating residues. This is
part of our job what we are doing.
Here, uh,
the presentation will be um,
divided into three parts. I will present
um,
introduction based
basic idea and applicability of our
software.
Then, uh, Maria will present how to run
and um, the examples of the analysis
with the use of Aqueduct will be
presented by Veronica.
Uh, so just
um, if you imagine the protein from
point of view of the small molecules of
the substrate or the inhibitor or any
other molecules, you can think that the
the molecule see the uh, surface which
is surrounding the the protein, but also
it could be just other um, um around.
So, it can see the surface or the
interior of the protein, yeah? Both of
them they are important, of course, from
our aspect. The question is how to how
to look into the second one. How to look
into the interior properties?
And there was a lot of um
uh tools developed
uh before based on
geometry um based approach, uh such as
Caver Mo or others, and they work very
nice, and they are very uh good for many
of things. They can detect the tunnels,
they can detect cavities. However, they
have some some some limitation.
One of the limitation is that since they
are using uh usually um spherical um
spheres to approximate the tunnels or
cavities, then uh you have problems with
facing uh and asymmetric features both
in case of the
interior cavities as
um tunnel entrance.
The second thing is that since they are
based on the mathematical algorithms,
they don't recognize the chemistry. So,
they
the same
um the tunnel with the same diameter uh
which will has the
charge or not charge residues are
treated in exactly in the same way. The
other thing is that even if you can
expand them to analyze molecular dynamic
simulation, then in a case of transient
tunnels,
they won't be detectable because you
need to see the pathways uh from the
beginning till end. Uh so, still from,
let's say, active site to the surface of
protein within one frame. If there is
some disconnection uh in a in a snapshot
of the MD simulation, then the tunnel is
not detected. So, we were thinking how
to um how to solve this problem, and one
of the uh our understanding was that if
we look into MD simulation, and this is
the frozen image from one snapshot, then
you can see that water molecules pretty
well penetrate also protein core. So,
then you can find them in the
uh inside
in active site in in in places
between so in the in the in the tunnel.
Um so, we had an idea to use this
possibility to search for the tunnels
and describe them. However, when we look
into single water molecule behavior,
this is 50 nanoseconds of simulation and
just single water molecule, you see that
the
trajectory of this molecule is pretty
pretty complex. So, in one moment we
were
facing the problem how to solve it, how
to approximate if we want really to use
this method since we have more than
sometimes more than
dozen thousand of of small molecules in
our system.
And how this is how we we develop the
Aqueducts, so the software to track and
to solve this this problem.
The basic the concept of the tool was
very
simplified.
So, um
Sorry.
So, if you
So, we are
not considering water molecules which
are out of the protein surface. Yeah, if
they are in the surrounding that we
neglect them completely.
Then we minimize
the search of the water molecules only
for those one which are in the object.
Object is something what we call the
place of our interest. So, for example,
let's say that this is active site
cavity.
Then you can see that we use some color
coding. So,
the red is
And then we are tracking only those
molecules water molecules or some small
molecules which were detected in in
object. So, in this case in active site.
And we are tracking them
how they are leaving and how they get
into the object or the active site. We
use the color coding so you see that
there is red pathways show
entering
molecules, blue
the leaving molecules, the green one
depict how the molecules behave within
the active site or the object of our
interest.
It can be also co-factors site or
whatever else and the yellow is the
other
possibility and that's the
molecules which are penetrating
interior, leaving object and coming back
to the object but they are not leaving
the protein. So it means they they are
just
going to other internal cavities and and
and and and and provide some description
of the of the interior.
The two other things which are important
it's are inlets. So inlets are the
points
locations in a in a space which
are on the border of the convex hole
which describe
the protein surface and they are just
informing us where the entry or the exit
of the tunnel is located.
How this concept look into
in practice?
That's the example of single water
molecule pathway
traveling through cytochrome P450 where
you can see that in one location was
entered then it was pretty long time
staying in one location because of the
gating residue which was closing
entry of this water molecule further.
Once the molecule the the residue side
chain moved then water molecule moved
forward it reached the active site a
heme and then leave by other tunnel
located in other place of the
of the protein.
Um this can you
analyze for all water molecules and then
you have statistical overview of all
information of all water molecules which
were traveling through
your protein, through your active site
during the entire simulation.
You can visualize this
on the left side you will see the shape
of the tunnels and you you see it's not
the circular, it's really representing
the dynamic of the
um of the protein and also the colors in
the identify the intensity of the number
of the inlets. So then you you see where
majority is um
is is is coming from or leaving.
And on the right side what you see, you
see again it's a cytochrome, you see the
picture of um
five, uh let's say um
tunnels, uh which I mean five entry
points to the protein um and and then
you you see the statistical interview
how many water molecules were detecting
traveling through particular entry.
And the
um the picture below the the circular
picture below, it's
maybe more complicated but it's very
very informative because it's providing
information about direction of the of
the
uh of the
water traveling through the
protein. So then you can say for example
that from 2B
uh tunnel through the entry of 2B, uh to
tunnel number four, the 30 um the
particular number 22% of water molecules
were were traveling. Yeah, so you you
you see the whole directions and you can
see the
uh the network of the tunnels, how they
it contribute to the travels of the of
the molecules.
That's a dynamic picture, and you can
also analyze the other features. So, you
can also analyze the local distribution
of
um of uh water molecules. For that, we
are using the the grid. So, then we are
calculating number of water molecules uh
or any other molecules uh during the
entire simulation um in a cubics. Uh
these cubics
um then provide statistics, and then you
can uh do several things. The one thing
is that you can allocate them uh
with the pathway, which usually for that
you uh simplify. And then, based on the
statistical distribution and using
Boltzmann equation, you can calculate
approximate the barrier of uh
um
the the barriers of the the the the
location of the um
most difficult um places to to move
uh water to transport water molecule,
then you can identify uh gating residues
or other obstacles which which um
contribute to the um
uh control of the of the transport of
the water molecules. And the second
thing is that uh you can also identify
uh the location which are the most uh
most often occupied by water molecules.
So, the density the local density of the
solvent is uh is uh the highest. And
that's we are calling the hotspots, and
they of course uh represent two uh
scenario. The one is where uh molecules
were trapped, and the other where they
strongly bind.
And um that's uh something what we can
use in a variety application.
So, uh
if you ask how we can use uh this this
type of analysis,
uh then the first is what I show already
is transport analysis. So then you can
depict the information about the
transportation of small molecules
through your protein. In many aspects,
you can identify also the very rare
events. So the pathways like here the
ones where only few molecules
go through and that can be used for
engineering. Then you can
use this information for structural
analysis of the of the proteins. This
will be one of the example of the
applicability and the
show by
Veronica.
You can a little bit understand
evolution of the protein especially in
if if the tunnels are important
due to the the protein uh
um catalytic machinery and
you can also use
it for drug design or inhibitors design.
Here specifically you can use not only
water molecules, but you can use also
a simulation in cosolvents and then you
can use every small molecules as a
specific molecular probe testing
different properties. Then you can build
pharmacophores
and based on on on hotspots and try to
design inhibitors which match the
pharmacophore. And
you can use also for protein engineering
and testing why particular mutations
cause difference in activity especially
where it it it reflects some
issue with transportation of the
substrate product release or also with
accessibility of the water. What is more
what is important also is that a product
is system independent. It's ligand
independent and it's highly flexible.
There is
several
models how you can use it to merge
different simulation together or to
or to divide single simulation into
counterparts just to see how the
dynamics is it's it's modifying the
transportation phenomena within protein.
And then
then
the the three examples which will be
shown later on
comes from our study one about epoxide
hydrolases
when we tested aqueduct on seven
different epoxide hydrolases and as you
can see here on this photo on the on the
slide
they were very diversified in case of
the how how network and how
accessibility of the access site it's
present. And you can see that the
mammalian has very
rich number of tunnels entries
where the the plants are preferential
one tunnel which is in main domain of
epoxide hydrolase and instead bacteria
has the entry the main entry and the
the primary entry through cup domain. So
it's kind of evolutionary
conserved
feature. And
this one of this
epoxide hydrolase from Solanum tuberosum
will be used as example
how to analyze the results to to
describe structure of of this of this
enzyme and it will be shown by Veronica.
The second example shown by Veronica
will be the work recent work of Agata
Raczyńska who unfortunately couldn't
join us today. And it was about
understanding of
substrate specificity of the pathways
for transportation of carbon dioxide
and
uh
formate through uh tunnels in
times independent formate dehydrogenase.
And also how
does it
work with the mutants and why this
mutant is
specific.
And the last one will be example about
drug design and here we again will go
back to epoxide hydrolase and this is
our study where we use several
cosolvents to map
hotspots in the interior and try to to
find the new
cavities, new
properties or new residues inside
protein interior which could be
targeted. The reason of this study was
that most I mean all of the inhibitors
uh
uh so far developed they didn't pass the
clinical trials so we were searching for
some inhibitors with unique features and
and it's will be also one of the example
on the last
last example. And that's I think all
from my side.
I would like only to share that uh
um most of the information will find of
our Aqueduct website dedicated for the
software and you can also find
um in the Living Journal of
Computational Molecular Science the
the tutorial which is very very complex
but it's going through all um
all steps of analysis and also of
preparation for the calculation which we
now
show in shorter version and it will be
shown by Maria Błufka.
Yes, so let's move to how to use
Aqueduct.
Next slide.
And
as you can see here this is like the
global picture how the Aqueduct software
looks like.
What we need for the Aqueduct to run the
calculations is
first of all the trajectory from our
molecular dynamics simulation together
with the topology file and also some
config file where we put all the
necessary information. What is important
here that Aqueduct use two drivers. The
first one called valve which is the main
driver to perform the whole analysis of
the trajectory the selected ligand in
the simulations and then the
complementary pond driver which does
this advanced analysis based on the
local distribution and local density
analysis that are to mentioned and this
driver pond uses all the information
that were previously computed by the
valve. So it means that for all the
distribution analysis that you want to
perform to obtain the information about
the density, the pockets, the hotspots,
first you need to track
your molecules of interest. We also have
a small application that helps to gather
and visualize the results and what also
Veronica show later is that the ultimate
output can be visualized in PyMol as a
PyMol session. So let's move further.
What is really important and what we
want to stress out here is like what are
the requirements
that you need to follow before running
the Aqueduct calculation. So first of
all what you need is the trajectory as I
said that contain information about both
the macromolecule of the interest here
usually the proteins and the small
molecules that are present
there. So usually water molecules but as
we can say it can be ligand independent
so we can track ion, substrate,
cosolvents, whatever you want and
whatever is present in your simulation
files.
And then before running the uh
calculations, you should really fix your
trajectory files by fixing the periodic
boundary conditions.
The protein the macromolecule of
interest should really stay in the
center of the box and all the movements
should be reduced.
Then uh
what is really important, especially for
the analysis of water molecules, it is
how frequent we save our
um
movements to capture them the movements
properly. So, we know that right now we
have longer and longer trajectories, but
we should really point out that the user
should aim for the
1 picosecond per 1 nanosecond ratio and
not exceed the 5 picosecond of sampling
because otherwise it would be really
hard to capture
um the relevant movements
of the molecules. And again, here
because like the concept of the Aqueduct
was strongly connected
with analysis of the water molecules,
please be aware of the proper protein
hydration before even running the MD
simulation. There are plenty of tools
that are available to place water
molecules properly inside the protein
cavities
or run sufficient equilibration of your
system to enable to
properly penetrate the solvent molecules
the
the protein interior.
And what is also important is that for
running the Aqueduct calculation, we
require a functional um Python
environment, preferably Python 3.9, and
we strongly suggest to do all the
Aqueduct calculation without the within
the virtual environment. Installation of
the Aqueduct is pretty straightforward.
It's described also on our website. So,
uh
it's really easy to
install the AQUEDUCT and then follow all
the
tutorials and run your calculations.
Okay, we can go further.
And here
what is important and what also Arthur
mentioned is like the definition of our
macromolecule and the object of the
interest. So, object is like the part
which interesting for us the most
because within this area that we define,
we track the the molecules.
Uh
and
usually it is uh
connected with some active site or
cavity within the protein molecules. So,
the definition
of that
uh
should be
should be done with uh
careful consideration and you should
really also define like which molecules
you want to track
that are passing this object of the
interest. And as Arthur said, we are
restricting our um
investigation area to the
scope, so the area within the analyzed
system
within the molecules should be traced.
We can go further.
And
for
you just to mention is that the AQUEDUCT
can be run in two different ways. For
guys, you can use either the graphical
user
interface which should um facilitate the
running the calculation and also
preparing the configuration file which
is necessary to run the calculation, but
you can also
do it just in the terminal and then just
by um
manually editing the configuration file.
For the new users, we strongly recommend
going through the graphical user
interface because then just
uh it guide you how to properly um
prepare your configuration file. We have
different levels of preparation of the
configuration file starting from like
very easy mode when we just trace uh
molecules of interest going through the
more advanced modes where you can also
perform more advanced analysis
such as hotspot analysis or pocket
analysis. And then we have also this
small graphical user interface called
Kraken for the visualization of our
results
where you just upload the
data that you obtained from the aquaduct
and you can obtain the um
uh nice visualization of your results.
And here what is really important and
what we want to stress out uh is the
proper definition of the
scope
because
by the defining or redefining the scope
uh we can
obtain like different information about
the trimming of the paths
and paths are important
because of the tracing of the molecule
and like area where the molecules
entered or uh leave the
macromolecule of interest. And the
general rule here is that the paths are
really limited to the scope area. So the
bigger the scope is, the longer are the
paths. So you can imagine that
if you analyze the protein, so if you
want to um
have your inlet so like the the
positions where the uh, molecules enter
or leave the
uh, scope to be located at the, uh,
greater distance, then you should define
the,
scope as the entire, uh, protein. And if
you want to go narrower and narrower,
you can, uh, go either through the
backbone definition or even the main C
alpha, uh, definition, uh, to obtain
shorter paths. And here is the example
on the, uh, left side you can see,
uh, how the,
um,
how the, uh, scope is defined as a, as a
protein. So, those, uh, purple lines
shows the definition and those small X's
show, uh, where the,
where the inlets, uh, will be located.
While with this, um,
uh, more deeper definition, uh, when the
scope is defined as a, uh, backbone,
the, the inlets, uh, will be, um, will
be displayed, uh, deeper,
uh, in a protein. We have also
additional, uh,
procedure for trimming the pathways, uh,
which was developed, uh, together with
the, um,
uh, Aqueduct. It's called, uh,
AutoBarber procedure. And,
uh, it, uh, works by trimming some, uh,
pathways accordingly.
Then, uh, as Arthur said, we can have
this, uh, information where the, the
molecules point out the entrance or the
exit area uh, to the interior of the
protein. And, basically, this is this
golden, uh, spheres area. We have many
such of the points. And usually, we want
to de-cluster them accordingly. Usually
based,
on the, uh, specific regions,
uh, where those molecules, uh, enter.
And there are, uh, many ways to divide
one big cluster into, uh, smaller
pieces. In Aqueduct, we implemented
different methods for that. And if you
go further, Arthur, you can see that
based on um
a choice that user is
that that the user is doing, we can
obtain like slightly different pictures
where the clusters will be located. So,
if we combine that with the information
from the paths, we can really
specifically cluster the inlets
accordingly to the entrance and exit
area.
And if we go further, we have this like
quantitative analysis that we can say
and Arthur already mentioned a bit about
that.
So, this is just like the graphical
visualization of the of the results
about the cluster size, the relative
flow between the
between the cluster, the intramolecular
flow, the directionality of the
of the flow, and also the
time within the simulation where
particular molecules entered
our scope and objects.
And if we go further,
we also have this information from the
distribution analysis when we can
analyze the
area
that were visited by our tracked
molecules. We can obtain information
about the the volumes, the space
which was penetrated
very often. This is the picture A where
the molecule spent
sufficient enough
of time. But, we also can obtain the
information about the maximal volume
that was penetrated by the
tracked water or other molecules.
And based on the density information, as
also Arthur said, we can obtain this
information about the hot spots. So, the
points within the macromolecules with
the highest local
density.
Uh
so, basically we can obtain the
information about specific location
where the molecules of interest were
attracted, trapped. We can see that they
differ um
in size here. So, the biggest molecules
this spheres means that the highest
density
was exactly in this specific area.
And
last but not least, as also Arthur said,
we can perform this analysis for the
um
not only one type of molecule
present in our system. If we have the
simulation with the cosolvents, and here
we define cosolvent as any um
um low molecular mass compound that can
be mixed with water molecules, and they
act as those like very specific probes.
And here we can clearly distinguish like
how different they can penetrate the
interior of the protein. For instance,
different parts than water molecules.
And it can give us like the insight
uh that we could
put particular focus on very specific
parts of the protein interior um based
on the type of the cosolvent or solvent
that we were tracking.
And with that we can go to the last
slide where we just show that how you
can visualize the results. It's very
simple but by
typing the
command to open the script for the
visualization and here in the in the
PyMOL you can see the loading of the
uh all the results when we have
information about the protein, the
clusters, and also all the paths between
uh the clusters and also like uh single
um paths of every single um
water molecule that was tracked.
Here. And with that we will go to
Veronica and practical examples.
Yes.
>> Yes. So, I hope you can see my screen.
Uh
as Arthur mentioned, we will go through
three different example of usage of
Aqueduct and we will focus on the
results visualized in PyMOL.
Uh for today we selected for you three
examples. So, simple water tracking,
then substrate tracking, and mutation
effect at and the end uh drug design
example. So, we will start from water
tracking. For this we use uh epoxide
hydrolase from potato. In this enzyme
water is very important for the reaction
of the hydrolysis. So, Maria introduced
you to the definition all of the things
we will see. So, I will not say that
once again, but uh here you can see uh
the scope. Scope is defined here as a
protein.
And you can see object. Object is
defined as a sphere
uh of the amino acid which create the
active site.
So, how you will visualize the results,
how you will look into your results,
it's up to you how you prefer to look on
the protein. Uh we prefer to look uh on
the subs- surface of the protein.
So, if you will visualize the surface
and you will open the cluster, you can
see where is your tunnels. So, here for
example, you can see one tunnel and you
can see that it's a hole. So, you are
sure that here should be some tunnel,
but if you will look, some tunnels can
be deeper a bit inside and they are not
visible on the surface because some
tunnels can be closed in the selected
frame. So, you can use the transparency
mode
on the surface and multi-layer is now
on. So, you will see the interior
cavities and your surface in one time
and this is how we prefer to look into
the session. So, here you can see that
in this epoxide hydrolase, we have one
main cavity
and we have the some tunnels here. The
third one is here and some inlets here
at from the other side.
So, based on analysis, you can see how
big your cluster are and you can see and
say that, "Okay,
this tunnel is the main tunnel because
the flow is
big here and the rest are used more less
frequently." You can also look into the
paths. Here you can see single row paths
for single water molecule. So, as Arthur
said, you can see that sometimes the
water is trapped in some parts. So, you
can analyze all of the places. You can
identify the amino acid which are
trapping the water, but you can also
look uh
on the path and
see where water enter to the protein
when leave protein. So, for example,
here you have the red color. So, it
means that the water enter by this small
two inlets cluster and it leaves the
protein through this one
bigger cluster.
And if you will look into the residues,
we are close here, you will see
that here we have as asparagine,
uh some glutamic acid and leucine. So,
they are responsible for this
opening of the small hole.
And you can also analyze here, for
example, the bigger tunnel. So, you will
see that here we have two isoleucines
which can work as a gate here, but also
we have two,
uh, aromatic residues, phenylalanine and
tyrosine, which are flexible and they
can,
uh, they move and they are responsible
for opening of this second smaller
tunnel. It's very worth to look into
your MD data to check which amino acid
are flexible and connect the results of
Aqueduct with the flexibility because
maybe you will find some important
gating residues and, uh, for example,
residues which you will, uh,
mutate in the future steps if you will
use, uh, the results for enzyme
engineering, for example. But, uh,
besides such a
single path look, look, you can look
into the results in the more more global
way. So, here you can visualize not only
a single paths, but the whole paths and,
for example, you can see, uh,
the paths between the tunnels. So, for
example, here you can see, uh, two four,
it means that the water enter by the
cluster number two and the leave protein
cluster number four. You can say see the
opposite side, so for cluster number
four to two.
And if you will open all of them,
you can see
that you have a lot of paths inside and
some are more dense, some less.
So, that's why how we look,
uh, for searching the density and
hotspots. But, when we move to the
hotspots, here are some uh, important
things.
You can also visualize paths of the
molecules which enter to the protein and
stay inside the end of simulation of the
paths with for the molecules which were
inside from the beginning of the
simulation and escape through different
tunnels.
Of course, you can see the hotspot, so
the most dense part
in the protein. And here, for example,
you can look into this red hotspot, and
you can see that it's close to the
catalytic histidine. So, it means that
somewhere here it's very important
water, very often, which is required for
the reaction of the hydrolysis.
And you can visualize
pockets. So, we are using mesh
visualization, and here you can see the
inner pocket, so the water was
most during the simulation was likely in
this cavity, and you can see outer, so
the whole space for the water, and you
can see the differences between the
pockets, so the inner is smaller than
outer, and this results can show you
very a lot of very useful information
because you can compare it with the
cavity for your crystal structure, for
example, and then you will see that the
cavity during the Aqueduct analysis is
much bigger, can try to find new spaces
in your protein, new cavities, and maybe
try to put new inhibitors in the new
space which you defined.
Each Aqueduct results give you a two
files which are not visualized in PyMOL,
so txt file and csv file.
And within this file, you can find a lot
of very useful statistic data. We will
go through this text file, so you can
see
first you can see all of your config
file configurations. So, you can find
how you define the object scope and so
on.
So, it's a lot of such a data. Then, you
can see how much frame you used in your
simulation, which type of molecules were
used to trace. And here you can see how
many water residues are found. So, it's
936.
And for this number of residue, we have
1,299
two paths. And the number of inlets is
2,532.
And you can think why we have more
inlets than
residues to trace. So, some residues can
enter the protein, then go out, go again
inside the protein, and so on. That's
why we have much more inlets than
residues to trace. Of course, you can
find all of the statistic about
clusters. So, the
clusterization
procedure, the cluster data, so how many
inlets you can find in your cluster, how
many incoming, outgoing, and different
type of statistic in the tables. They
are all described.
Here is one big table. So, here in this
table, you can
see a lot very useful statistic for
single water molecule based on ID of
this molecule. You can even found the
time, so in the frame when the molecule
entered the active site or entered the
scope. And when molecule leaves
from the active site. So, we will go to
the second example. Here we have formate
dehydrogenase is the enzyme that perform
the reaction of reduce CO2 to formate
and within this example we trace the
substrate so we trace CO2 we also trace
formate and we prepare some mutation. So
here you can see as a mesh visualize the
tunnels detected by Caver. You can see
that Caver detect two very narrow
tunnels and if you will look into the
paths you can see that for example CO2
it's using two tunnels. Here the smaller
tunnel and the big
tunnel which is here and what is
important here you can still see here
the Caver tunnel is very narrow but CO2
use much bigger cavity for the transport
to the active site.
And if we will just look into single
paths
you can see that some example of enter
of CO2 to the active site and escape
from the one big cluster but if we will
go to different paths for example here
you can see that this tunnel probably
were closed in that time. So CO2
molecule decide to just go back the same
tunnel and escape and what is very
curious here that this tunnel is very
narrow. Here you can even see the
space between the cavities.
So if you will look into amino acid in
this area you will see that here we have
a gating residue is tryptophan and is
one arginine residue
and we decided to mutate this residues.
So if you will look into the mutation
here is glutamic acid and you will see
that glutamic acid can create of
arginine the salt bridge.
So if you will look how CO2 is going
through the protein within this
mutation, so we will see that whole
transport is almost by this one cluster.
The second cluster is used into the
smaller way. So, it's just two inlets
here.
And uh here we can also look into the
single path and we will see that here,
close to this salt bridge, water was
trapped CO2 was trapped for some part of
simulation and then escape when this
bridge disappears. So, even such a
change like small bridge, which is not
so strong interaction, can change the
flow of the substrate or the solvent
within the protein. And if we will go to
the
uh formate,
we can see that
uh here we have also these two tunnels
detected by Caver.
We will see that formate is using just
the one main tunnel to go uh to the
active site of the protein probably
because of uh the size of the formate.
Formate is bigger than CO2 and this
second tunnel is just too narrow to be
used
uh for
enter or leaving the molecule.
And uh
the last example
it's uh example of human soluble epoxide
hydrolase. Here you can see the
structure of human soluble epoxide
hydrolase with all of the crystallized
uh inhibitors. So, this enzyme performed
the hydrolysis reaction of very
important epoxyeicosatrienoic acid,
which can be used as uh potential
non-inflammatory compounds in the
treatment of many diseases.
And for today, we don't have any good
drug on the market, which is working
actually no drug pass
uh clinical trials.
So, here we used Aqueduct
with cosolvents. So, you can see that we
have different type of cosolvents and we
can describe the pharmacophore features
to each of the molecules. So, we have
acetonitrile, which is the hydrogen bond
donor and hydrophobic compound. We have
dimethyl sulfoxide, which is hydrogen
bond acceptor. It's here.
Green one, we have methanol. Methanol is
hydrogen bond donor. Phenol,
which is also aromatic. And here it's
not so good visible. Phenol, it's here.
So, small hotspots.
Then we have urea, which is the bond
donor and acceptor. And of course,
water, which is also hydrogen bond donor
and acceptor. And what is curious here
that if you will look into detail into
the inhibitors, you can see that if you
have a group which are able to create
hydrogen bonds,
so they have nitrogen or oxygen, water
hotspots are very close. For example,
here we have more hydrophobic part and
here we can see the hydrophobic part of
inhibitors. Here where is this
phenol, you can see a lot of aromatic
strings
within the inhibitors. So,
it's kind of the proof of
our concept that we can see
the description of the places based on
the possible interaction within the
protein. And now if we will switch off
the inhibitors, we can see that we have
a very nice map of potential
interactions or even the description of
the places. So, we can say a lot
which for example part of the protein of
the cavities more hydrophobic,
hydrophilic, which prefer to have some
aromatics
part of the inhibitors. So, based on
that, we can design the inhibitors or
try to even find by virtual screening
the compounds which will fit to our
cavity based on the interaction which we
can
describe here. And now we are working to
go more into the detail within this. So,
we are trying to extend Aqueduct a
little bit to describe better the
properties of the cosolvents
and to automatize the
selection of the features
for some drug design project. So,
I think that's all from my side and I
think that we still have a time for some
questions.
So, if there are
Yes.
>> one. Yeah.
Yes, there was one.
I lost the window now. Uh sorry, there
was one regarding membrane proteins. So,
I think Maria, you can
elaborate about that. The question was
if you can use Aqueduct for membrane
protein and if there are any limitation
for that.
Um I mean, yeah, basically if you
have such a big system like a membrane
protein, um I suppose like I don't like
GPCRs in certain
inserted in some membrane like a
transmembrane protein like a receptors.
Obviously, you can use Aqueduct for this
analysis. Like the way here is like the
definition of again like the scope and
the object of the of the interest. But
and I think that that was something
about the ion channels that uh
you can also use that for the for the
ions channels. Here like there is no
problem with the definition
itself.
But the issue here is that uh
you
should have like uh
If you have a very
hydro
uh
hydrophilic
environment and when it's really hard
for instance for water molecules to uh
pass,
then you will have like very rare
sampling in your MD simulation and then
maybe it can be hard to
um
to describe like the particular
events uh
present in your system. But for instance
for the ion channels you can
track the um
ions in in the channel and here it can
be defined as a molecule to track and
you can obtain all the information about
the what
the specific kind of ion is doing within
your channel.
Okay, there's more questions here.
Thank you very much for the lecture and
the cool results. I'm curious
how long does the simulation and the
analysis itself take?
The simulation by itself it's just MD
simulation so and the the aqueduct
depends I mean
calculation of aqueduct depends on how
many water molecules in fact you
tracked. So of course if you have a
larger number then it's getting longer
and longer and this is something what is
difficult to
um to uh
to to paralyze. So, so we have some kind
of problems if the system is too large
or too long. Uh it it it can be a
problem. Uh we have some kind of
benchmark
uh in
publication
which was done for 50 or 100 nanoseconds
if I remember and number of different
water molecules.
Um
what we truly recommend when you are
running analysis is to
pick um
uh some smaller uh
time uh or part of the simulation first
just to really design properly the scope
and the object uh just to test what
settings are the best to answer your
questions. And then truly to think twice
um
how to minimize the effort for
calculation because the number of water
I mean, once you are reaching really
high number of of
of molecules to track,
then it starts to be long-term
project. The largest system which we
used was Maria in fact using it was TLR
um
part of the TLR receptor together with
other protein and then we define the
object in on the interface between these
two proteins to
to track the water molecules which
participate in reaction of the of the
loop cleavage. So, you can use very high
system, but if you properly define the
object and the scope, then you can
minimize the cost of the calculation.
You need to be really careful and smart.
Exactly. Here, like you don't need to
define the scope as an entire complex.
You can focus only on the specific part
and that would be something that we
recommend to overcome the limitation
with the number of atoms in your system.
And sometimes it's better to use
shorter trajectories
in replicates than
one very very
long one.
for our calculations.
Okay.
Then next one. Oh, you want to continue?
Yeah, what are the limitations? So, the
limitation is
it's it's really the number of I mean,
we didn't test
what is the maximum number because it's
really system dependent also the
resources dependent.
So, it's it's hard for me to answer this
question. Basically, what you can do
always you can treat system in a little
bit different way and just divide
simulation into small party small parts
smaller parts if this very very large
number of track molecules and then you
can solve the issue.
Regarding the updates, one thing which
we are thinking
and we are working on it's it's it's
it's it's for pharmacal for design but
this is kind of complicated problem and
uh
and and we are trying to solve it
already 2 years. So,
once we will solve for sure we will
share to
to community.
Then there was a question regarding
ramp and ligand exits and compare of uh
Yeah, so this is this is possible.
I mean,
the power of Aqueduct is in statistic.
Yeah, so if you have just a simple
single molecule which you are tracking,
then Aqueduct is not the best tool for
that. Yeah, but if you want to to
how water is assisting your substrate,
then of course you can do comparative
analysis, so you can run a product for
one and the second event and try to
to compare this data. We also built some
feature which was tested
on on one of the example
exactly trying to calculate the barrier
of the entry of the substrate, so then
we were what we were doing we were in
fact merging fragments from simulation
from different simulations, so let's say
that you have 10 simulation when
some event occurs very rarely, but you
pick only the the parts of the
simulation when the event took place and
emerge them together and analyze
together. So this is also somehow
possible one of the component, however,
it is tricky and I think the most
complicated
job for aquaduct to to do. But for sure
if you have long enough sampling for
both states, so ligand bound and ligand
dissociated, then you can do comparative
study and for sure you will see
difference in opening or closing if they
are yeah or some conformational changes
which will
differentiate the the for example
interior space.
uh
300,000
We have time for one more question, so
maybe you can read one aloud and then
someone quickly and then
I will have to close.
um
So regarding the last question, if there
is any problem with solutes like
glycerol, glucose or another
another
class
um again
it depends on how many uh, molecules you
have. There is no problem for Aqueduct
regarding what it is to track. Just to
to be to have consistent or good
results, you need to have simply more
than let's say 20 20 molecules in a
system. Yeah, if you have less than 20,
then statistic will be very very small.
Or you need to have very long
simulation
multiple replicas.
All right. Any other questions, you can
please put them in the BioExcel
forum and
the presenters will try to answer them
the coming week. Yes, we will.
Yeah, thank you for opportunity.
>> I'll just take my this opportunity to
tell about the upcoming webinars. There
will be one in still at the end of this
month
about DynaPyn
and then the next one is actually in
April
14th, multiscale simulation of
biomembranes.
So, thank you to everyone.
Thank you.
>> Thank you.