Video summary
The second quarter update for OpenFF highlights significant progress across science and sports fields, despite the team dedicating considerable time to workshops and conferences. The primary focus remains on the water model co-optimization project, which is advancing along three main directions: developing a minimum viable product (MVP) using force balance troubleshooting, creating a new PyTorch-based architecture for future optimizations, and refining data workflows. While progress has been slightly slower due to competing compute needs with other major projects, benchmarking efforts have begun on the FrontierFF simulation performance. Additionally, discussions continue with NIST regarding an expanded dataset release, while specific issues like poor performance in alcohol-water mixtures are being addressed through electrostatics rescaling and grid searches for optimal scaling factors.
A key development this quarter is the creation of a new library designed to replace the OpenFF evaluator by managing computer operations without relying on force balance or co-optimization terms. This tool, built on a PyTorch approach, has been ported with existing functionality for handling machine learning datasets and scaled up from prototypes in the Descent library to production levels. Beyond this technical infrastructure update, new releases of Descent version 0.6 have introduced support for regularization during training and periodic systems when calculating energies, alongside improved documentation. Furthermore, Danny and Finlay have released Presto as a mature product featuring a bespoke fitting workflow that utilizes SME and Descent to train force terms on machine learning potential data sets.
The update also details substantial work on Stage 2.4 projects involving the recently released B-Dance dataset, which offers optimized geometries and Hess data orders of magnitude larger than previous archives. Analysis of this dataset has revealed parameters with bimodal distributions that may require splitting for specific chemistries to improve accuracy. Simultaneously, efforts are underway to augment benchmark sets to cover sparse chemistries more effectively. The protein force field project continues its dual-path strategy: one approach involves patching problematic parameters in the Rosemary Alpha model back to Stage 2.1 values to fix errors while maintaining performance, and a second path explores integrating Stage 2.3 typing with minimal perturbation of torsions. Initial results from the patched version show mixed outcomes on specific benchmarks like helicity and NMR observables, prompting further investigation into sampling differences versus meaningful parameter changes.
Finally, significant strides were made in membrane force field development led by Julianne, who aims to create parameters performing well for both small molecules and membranes. After identifying issues with previous Stage 2.2 models regarding order parameters and relaxation times, a new workflow utilizing Sage 2.3 significantly improved performance on metrics like area per lipid and bilayer thickness, bringing results in line with existing force fields like CHARMM. Although challenges remain with form factor quality and halogenated alkane properties, retraining van der Waals terms to an alkane dataset yielded promising improvements in relaxation time. Looking ahead, the team plans to continue refining the Rosemary Alpha model for release, expand collaborations on standardized benchmarks through OpenFE, advance Stage 2.4 fitting efforts, and develop a liquid force field, with updates expected next quarter.
Read the full video transcript
Hi everyone, Lily here. This is the
second quarter update for science and
sports fields at open FF. A bit of a
shorter one than Q1 as most of the team
spent a lot of time variously preparing
for and attending the Mayf workshop,
attending other conferences or often PTO
sik or rental.
That said, we've still managed to make a
fair bit of progress on our general road
mapap deliverables. So without further
ado, let's get into it.
Last time we talked a lot about our
water model coopization project and the
various updates we had made to software
and data workflows. Broadly this can
still be divided into three major
directions. Firstly fitting a
three-sight uh MVP using force balance
troubleshooting particular alcohol and
aiming properties and creating a new dim
sim library for future optimizations
using um a new pytorch based
architecture.
The MVP part of this is still in
progress. Um unfortunately progress has
been a bit slower on this than we would
like partly due to our GP compute needs
starting to conflict with other major
projects that we've been working on.
So we are still pre-coribrating our
expanded public data set using Sage 2.3.
Um however Chris has started
benchmarking and optimization simulation
performance of FrontierF. So this
bottleneck will hopefully relax soon.
We're also still working on discussions
with NIST to release an expanded data
set with more properties that we can
train and validate to.
Uh Chris has also been looking more at
particular subsets of properties that
perform poorly focusing initially on
alcohol and water mixtures particularly
on NWS mixing. Uh we we think that this
may be an electrostatics issue and
previous work has found that rescaling
alcohol charges substantially improved
alcohol properties.
So working off a base of our ashc
charges, Chris has been performing grid
searches over scaling factors and
identified a set of values that perform
well for simple alcohols.
He's now evaluating the transferability
of these scaling factors to mixtures of
alcohols and dials and is starting to
look into whether optimizing the bandw
parameters with the new charters will
give further improvement. Once we've
settled on a good set of parameters, we
will need to work out how to best for
this into our existing chart model.
Whether we can do this um as a
post-processing step or independent step
or whether we'll need to integrate this
into a DC itself.
Lastly, matters continue working on
developing a new library for computer
management so that we can move on from
force balance and co-optimize veance and
vanderals terms using a pytor based
approach.
This library can be best thought of as a
replacement for open ff evaluator and
over the past couple months he's ported
over existing functionality for handling
the ML data sets and started working on
scaling up the approach prototyped in
the descent library to apply to
production level fifths. This has meant
both extending the property types um
already implemented in descent as well
as experimenting with and converging on
an approach to compute management uh
with prototyping on natural clusters.
Last but not least, we have a couple
other updates. So firstly, since last
quarter, we've made a new release of
descent uh version 0.6 that adds support
for regularization during training and
add support for periodic systems uh when
calculating energies. We've also tried
to make descent more user friendly by
adding more documentation. Secondly,
just highlighting that uh Danny and
Finlay have released Presto as a mature
product with a pre-print on Chem
Archive. So, please do check it out if
you're interested. Uh this is a bespoke
fitting workflow um that uses uh SME and
descent uh to to train for your terms to
MLP data.
Moving on to stage 2.4 work. Uh this is
fairly short but we have been looking at
the recently released data set from by
dance. So this data set also called them
contains optimized geometries, torsion
scans and hess data [clears throat] at
our level of theory and is orders of
magnitude larger than what we have in QC
archive. So it seems like a very
promising source for additional training
and benchmarking data. We've already
begun using it in some protein for fits
which I'll talk a bit about later. But
I've also started looking at whether it
can inform future stage fits as well.
For example, we've identified uh a
number of parameters that have biodal or
multimodal distributions of predicted
equilibrium values uh suggesting that
they may benefit from being split to be
more specific as to which chemistries
they cover. We're still in the progress
of looking through the data set and
planning the S the next stage 2.4 refit.
In terms of key next steps, we also want
to augment our existing benchmark sets
to increase coverage of sparse
chemistries and ensure all our
parameters are covered.
Uh all right on to the protein fossil
project.
Just to recap, we started this project
several years ago with a goal of
co-optimizing parameters to small
molecule and protein data for a combined
force that performs well on both small
molecules and proteins. Uh and this is
being led and driven by chapter from the
GSA lab. As an initial target, we're
aiming for similar accuracy and
performance to existing protein force
fields as benchmarked across four tiers
of benchmarks that increase in
computational expense and structural
diversity.
What we have tended to find is that with
our force field candidates, we do quite
well on tier one benchmarks, which
include QM targets and scale couplings
of short and structured peptides. But
our candidates have difficulty with
benchmark targets that assess secondary
structure, such as keeping proteins
folded in long time scale simulations.
Last year we released a rosemary alpha
force field that was fed to small
molecule and peptide QM data as well as
to the GB3 and helical peptide targets
or the tier 2 benchmarks.
On benchmarks rosemary alpha achieved
good performance on most of the targets
in tier 2. Although performance on head
egg white was still worse than our 14
SB. Last quarter we analyzed tier 3
benchmarks that indicated this was a
general force field problem but further
refitting to hen egg white lysazes
targets didn't improve general
performance.
In addition, we did some small molecular
benchmarking and found that rosemary
alpha introduced some pathologies mostly
around sulfur, phosphorus and caroxilate
chemistries.
We did decide to move rosemary alpha to
release as a peptide force field once we
fixed these while continuing to work on
folder protein behavior. But Chapen
found that doing a normal small molecule
and peptide QM refit from rosemary alpha
degraded performance on tier 2
benchmarks even when he froze the six
specific protein torsions that you tune
to the tier 2 targets.
So instead we decided to move forward in
two parallel efforts over this quarter.
Chapen looked at the easiest path
forward which was just to patch uh the
problematic parameters in rosemary alpha
to the stage 2.1 values for that refit.
This is the smallest modification we can
make to rosemary alpha. So it is the
most likely to keep performance on
proteins where it currently is. In
parallel, open FF would move along this
pathway where we would take the rosemary
alpha and integrated with stage 2.3
typing. We will refit veance parameters
while perturbing torsions as little as
possible and if this retains protein
performance status rosemary 3.1.
Oops.
Uh both tasks are still in progress, but
here is the current status. In pass one,
which is just patching the problematic
parameters, Chabin identified eight
bonds and two angles as the key targets
to fix. He made a new force field called
rosemary alpha 1B where these are
patched back to the stage 2.1 values.
This force field successfully fixes the
large errors in problematic parameters
while giving the same performance and
overall Q metrics as stage 2.1.
Next, he benchmarked this force field on
rosemary targets. On tier one targets,
this forio performs very similarly to
the original rosemary alpha. Uh on tier
2, however, we see some differences.
Firstly, rosemary alpha 1B results in
higher holicity on the AQA3 benchmark
than rosemary alpha. In this particular
benchmark, none of the parameters
changed affected these simulations. So
this is wholly due to sampling
differences between the replicates.
Secondly, rosemary alpha 1B unfolds the
GB3 alpha helix and has higher error on
NMR observables uh than rosemary alpha.
For GB3, there was one parameter that
changed between rosemary alpha and
rosemary alpha 1B, but this is just a
minor modification of a caroxilicate
bond. So, we're still trying to work out
if this is a meaningful difference or
more differences in sampling.
In terms of next steps, we're focusing
on benchmarking this force more. We will
proceed with the tier three benchmarks
of rosemary alpha 1B starting with
protein binding for energies. We will
also run an additional three replicates
each of rosemary alpha and rosemary
alpha 1B to uh on the GP3 target to
further characterize differences between
the performance on the two force fields.
For the second path uh this involves a
greater change to the force field. So we
did some more experimentation. Our first
step was to retype parameters from stage
2.1 to use the stage 2.3 partitioning
although we left the vanderal parameter
values untouched to retain protein
behavior.
We then investigated the impact of a
number of variables. The first was data
which data set we used for deriving
initial bond and angle values via the
modified simaria method as well as which
data sets we use for training. Secondly,
torsional flexibility. Uh whether we
completely froze all torsions or refit
them using prior widths of 0.1, one or 5
kilo calories per mole. Lastly, we also
experimented with uh different
weightings between training targets.
In general, we found that the
combination of variables that gave us
the best performance was to use initial
values calculated from larger by dance
or theme data set and to train to a
combination of QC archive and theme
data. In addition, using the default
waiting of the different training
targets gave us the best results. We did
observe that freezing torsions resulted
in worse performance on bonds than
allowing them to optimize. So we are
still experimenting with allowing
torsions to optimize with an arrow
prior.
We found that using a pri width of 1
kiloalorie per mole is enough to improve
performance on bonds and tier 2
benchmarks is still ongoing to see if
this will negatively impact uh folder
protein performance.
Lastly, Ashley is still continuing work
on benchmarking the rosemary alpha force
field on cyclic peptides. This quarter,
the simulations have completed and she
has started analyzing the results which
we'll hopefully present next quarter.
Um, on deliver, there have been a few
updates here.
As a reminder, this project is being
driven by Julianne from the shirts lab
and the goals are fairly similar to the
protein force field. getting parameters
that perform well on both small
molecules and membranes and a force
field that performs in line with
existing membrane force fields.
As a means of measuring performance,
Julianne has two sets of benchmarks that
she's using to compare performance. The
first is alkane benchmarks uh such as
torsion energetic profiles uh physical
properties and transcouch ratios. The
second incorporates actual membranes and
measures common targets such as area per
lipid blayer thickness order parameters
uh form factors and she's also been
looking at relaxation time.
As with the protein force field, I'll
just quickly summarize where we were at
the start of the quarter. So on
benchmarks of previous stage force
fields, primarily stage 2.2, Julian
found that they performed in line with
other lipid force fields on metrics like
area per lipid and blayer thickness. Um
however they performed pretty badly on
order parameter quality and were quite
slow to relax. She tried a number of
different refits and found that
retraining to a combined small molecule
and lipid data set uh didn't shift the
torsional terms for alkan profiles which
as you can see here get the relative
energies of the trans and gash component
is pretty incorrect.
Retraining to an alkan only data set
improved this particular target. Uh but
actually matching the Q profile exactly
resulted in overestimating the trans
gash ratio in simulation.
She also ran into repeated issues with
training phosphate parameters to QM
which may be because the constraint
optimizations that we use uh during
training with force balance result in
very different confidence when optimized
with mm relative to QM.
The workflow that she saw the most
success on was to retrain veance terms
to an alkane QM data set and then
retrain bands parameters to a combined
alkan and methylacetate physical
property data set. This improved order
parameters although not to the same
extent as of the lipid force fields.
This quarter, following the full release
of Sage 2.3, she reran her workflow and
benchmarks using Sage 2.3 as a starting
point. Interestingly, Sage 2.3 by itself
already improves massively on Sage 2.2
with order parameters and is now
performing similarly to other lipid
force fields like charm. Form factor
quality is still low, though suggesting
possible issues with head group
parameters, and the relaxation time is
still pretty high.
Retraining to arcane data from stage 2.3
and porting the refitted vanderal terms
from the sage 2.2 fit got the stage 2.3
veance one force field. This improves
relaxation time but decreases form
factor quality. Interestingly this
source field also decreases performance
on uh general physical properties
especially halogenated alkanes.
Julian also tried briefing phosphate
parameters using the presto library from
the coal group which was a lot more
successful in matching QM profiles than
force balance workflows.
That being said, performance on limited
benchmarks definely shift substantially
using these new terms um including in
areas of concern.
In terms of next steps, Juliana will
keep benchmarking her refit force build
on a wider set of liberties.
To address some of the issues with
overfitting to QM and physical property
performance, she's planning to try
copying veance of wild parameters as
well as increasing the size of the
training sets for her fits.
Lastly, on to QM data sets and
collaborations, both of which are being
masterminded by Jennifer Clark.
Over the past quarter, we have added a
few more data sets to the QC archive.
This includes a subset of molecules from
SPICE 2 that we identified as
problematic for various reasons, such as
overlapping atoms. This is only about
2,300 compliments, but we thought it
would be useful to fix for force field
training and validation workflows. We
also added a small data set covering PEG
fragments as part of a training tutorial
for fitting a force field for polymer.
And lastly, the initial architecture
data set for diverse organometallic lian
complex structures has completed and Jen
is in the process of setting up larger
one for the ongoing genetic and cader
lab collaboration.
Speaking of collaborations, open FF has
been collaborating with open FE to build
a standardized benchmarks repo. I
mentioned this last time as well, but
just as a reminder, this repo supports
relative binding for energies and
absolute salvation for energies. It
includes code for discovering uh
available benchmark sets and includes
standardized sets of partial charges to
avoid charge variation affecting
results. The repo also allows both
building new networks and using
previously generated ones and includes
comprehensive samples. We're pretty
excited about this project as it'll make
it much easier for Open FF to run our
benchmarks in the future. And indeed,
Jen is actually the one who is going to
be running the tier three RBFs for the
protein for project.
At this point, this collaboration is
largely complete. There are still some
odds and ends to wrap up, but most of
the minimum goals have been hit.
Our main focuses for the next quarter
are to continue with our ongoing work,
focusing on the water model refit,
getting the Rosemary alpha closer to
release, and continuing collaborations.
We will also keep working towards stage
2.4 and helping with development of a
liquid force field. So, we'll hopefully
update on this project six quarter. See
you next time.