Submind YouTube summaries
Thumbnail for July 2026 Science Update

July 2026 Science Update

Watch on YouTube

Video summary

The second quarter update for OpenFF highlights significant progress across science and sports fields, despite the team dedicating considerable time to workshops and conferences. The primary focus remains on the water model co-optimization project, which is advancing along three main directions: developing a minimum viable product (MVP) using force balance troubleshooting, creating a new PyTorch-based architecture for future optimizations, and refining data workflows. While progress has been slightly slower due to competing compute needs with other major projects, benchmarking efforts have begun on the FrontierFF simulation performance. Additionally, discussions continue with NIST regarding an expanded dataset release, while specific issues like poor performance in alcohol-water mixtures are being addressed through electrostatics rescaling and grid searches for optimal scaling factors. A key development this quarter is the creation of a new library designed to replace the OpenFF evaluator by managing computer operations without relying on force balance or co-optimization terms. This tool, built on a PyTorch approach, has been ported with existing functionality for handling machine learning datasets and scaled up from prototypes in the Descent library to production levels. Beyond this technical infrastructure update, new releases of Descent version 0.6 have introduced support for regularization during training and periodic systems when calculating energies, alongside improved documentation. Furthermore, Danny and Finlay have released Presto as a mature product featuring a bespoke fitting workflow that utilizes SME and Descent to train force terms on machine learning potential data sets. The update also details substantial work on Stage 2.4 projects involving the recently released B-Dance dataset, which offers optimized geometries and Hess data orders of magnitude larger than previous archives. Analysis of this dataset has revealed parameters with bimodal distributions that may require splitting for specific chemistries to improve accuracy. Simultaneously, efforts are underway to augment benchmark sets to cover sparse chemistries more effectively. The protein force field project continues its dual-path strategy: one approach involves patching problematic parameters in the Rosemary Alpha model back to Stage 2.1 values to fix errors while maintaining performance, and a second path explores integrating Stage 2.3 typing with minimal perturbation of torsions. Initial results from the patched version show mixed outcomes on specific benchmarks like helicity and NMR observables, prompting further investigation into sampling differences versus meaningful parameter changes. Finally, significant strides were made in membrane force field development led by Julianne, who aims to create parameters performing well for both small molecules and membranes. After identifying issues with previous Stage 2.2 models regarding order parameters and relaxation times, a new workflow utilizing Sage 2.3 significantly improved performance on metrics like area per lipid and bilayer thickness, bringing results in line with existing force fields like CHARMM. Although challenges remain with form factor quality and halogenated alkane properties, retraining van der Waals terms to an alkane dataset yielded promising improvements in relaxation time. Looking ahead, the team plans to continue refining the Rosemary Alpha model for release, expand collaborations on standardized benchmarks through OpenFE, advance Stage 2.4 fitting efforts, and develop a liquid force field, with updates expected next quarter.
Read the full video transcript
Hi everyone, Lily here. This is the second quarter update for science and sports fields at open FF. A bit of a shorter one than Q1 as most of the team spent a lot of time variously preparing for and attending the Mayf workshop, attending other conferences or often PTO sik or rental. That said, we've still managed to make a fair bit of progress on our general road mapap deliverables. So without further ado, let's get into it. Last time we talked a lot about our water model coopization project and the various updates we had made to software and data workflows. Broadly this can still be divided into three major directions. Firstly fitting a three-sight uh MVP using force balance troubleshooting particular alcohol and aiming properties and creating a new dim sim library for future optimizations using um a new pytorch based architecture. The MVP part of this is still in progress. Um unfortunately progress has been a bit slower on this than we would like partly due to our GP compute needs starting to conflict with other major projects that we've been working on. So we are still pre-coribrating our expanded public data set using Sage 2.3. Um however Chris has started benchmarking and optimization simulation performance of FrontierF. So this bottleneck will hopefully relax soon. We're also still working on discussions with NIST to release an expanded data set with more properties that we can train and validate to. Uh Chris has also been looking more at particular subsets of properties that perform poorly focusing initially on alcohol and water mixtures particularly on NWS mixing. Uh we we think that this may be an electrostatics issue and previous work has found that rescaling alcohol charges substantially improved alcohol properties. So working off a base of our ashc charges, Chris has been performing grid searches over scaling factors and identified a set of values that perform well for simple alcohols. He's now evaluating the transferability of these scaling factors to mixtures of alcohols and dials and is starting to look into whether optimizing the bandw parameters with the new charters will give further improvement. Once we've settled on a good set of parameters, we will need to work out how to best for this into our existing chart model. Whether we can do this um as a post-processing step or independent step or whether we'll need to integrate this into a DC itself. Lastly, matters continue working on developing a new library for computer management so that we can move on from force balance and co-optimize veance and vanderals terms using a pytor based approach. This library can be best thought of as a replacement for open ff evaluator and over the past couple months he's ported over existing functionality for handling the ML data sets and started working on scaling up the approach prototyped in the descent library to apply to production level fifths. This has meant both extending the property types um already implemented in descent as well as experimenting with and converging on an approach to compute management uh with prototyping on natural clusters. Last but not least, we have a couple other updates. So firstly, since last quarter, we've made a new release of descent uh version 0.6 that adds support for regularization during training and add support for periodic systems uh when calculating energies. We've also tried to make descent more user friendly by adding more documentation. Secondly, just highlighting that uh Danny and Finlay have released Presto as a mature product with a pre-print on Chem Archive. So, please do check it out if you're interested. Uh this is a bespoke fitting workflow um that uses uh SME and descent uh to to train for your terms to MLP data. Moving on to stage 2.4 work. Uh this is fairly short but we have been looking at the recently released data set from by dance. So this data set also called them contains optimized geometries, torsion scans and hess data [clears throat] at our level of theory and is orders of magnitude larger than what we have in QC archive. So it seems like a very promising source for additional training and benchmarking data. We've already begun using it in some protein for fits which I'll talk a bit about later. But I've also started looking at whether it can inform future stage fits as well. For example, we've identified uh a number of parameters that have biodal or multimodal distributions of predicted equilibrium values uh suggesting that they may benefit from being split to be more specific as to which chemistries they cover. We're still in the progress of looking through the data set and planning the S the next stage 2.4 refit. In terms of key next steps, we also want to augment our existing benchmark sets to increase coverage of sparse chemistries and ensure all our parameters are covered. Uh all right on to the protein fossil project. Just to recap, we started this project several years ago with a goal of co-optimizing parameters to small molecule and protein data for a combined force that performs well on both small molecules and proteins. Uh and this is being led and driven by chapter from the GSA lab. As an initial target, we're aiming for similar accuracy and performance to existing protein force fields as benchmarked across four tiers of benchmarks that increase in computational expense and structural diversity. What we have tended to find is that with our force field candidates, we do quite well on tier one benchmarks, which include QM targets and scale couplings of short and structured peptides. But our candidates have difficulty with benchmark targets that assess secondary structure, such as keeping proteins folded in long time scale simulations. Last year we released a rosemary alpha force field that was fed to small molecule and peptide QM data as well as to the GB3 and helical peptide targets or the tier 2 benchmarks. On benchmarks rosemary alpha achieved good performance on most of the targets in tier 2. Although performance on head egg white was still worse than our 14 SB. Last quarter we analyzed tier 3 benchmarks that indicated this was a general force field problem but further refitting to hen egg white lysazes targets didn't improve general performance. In addition, we did some small molecular benchmarking and found that rosemary alpha introduced some pathologies mostly around sulfur, phosphorus and caroxilate chemistries. We did decide to move rosemary alpha to release as a peptide force field once we fixed these while continuing to work on folder protein behavior. But Chapen found that doing a normal small molecule and peptide QM refit from rosemary alpha degraded performance on tier 2 benchmarks even when he froze the six specific protein torsions that you tune to the tier 2 targets. So instead we decided to move forward in two parallel efforts over this quarter. Chapen looked at the easiest path forward which was just to patch uh the problematic parameters in rosemary alpha to the stage 2.1 values for that refit. This is the smallest modification we can make to rosemary alpha. So it is the most likely to keep performance on proteins where it currently is. In parallel, open FF would move along this pathway where we would take the rosemary alpha and integrated with stage 2.3 typing. We will refit veance parameters while perturbing torsions as little as possible and if this retains protein performance status rosemary 3.1. Oops. Uh both tasks are still in progress, but here is the current status. In pass one, which is just patching the problematic parameters, Chabin identified eight bonds and two angles as the key targets to fix. He made a new force field called rosemary alpha 1B where these are patched back to the stage 2.1 values. This force field successfully fixes the large errors in problematic parameters while giving the same performance and overall Q metrics as stage 2.1. Next, he benchmarked this force field on rosemary targets. On tier one targets, this forio performs very similarly to the original rosemary alpha. Uh on tier 2, however, we see some differences. Firstly, rosemary alpha 1B results in higher holicity on the AQA3 benchmark than rosemary alpha. In this particular benchmark, none of the parameters changed affected these simulations. So this is wholly due to sampling differences between the replicates. Secondly, rosemary alpha 1B unfolds the GB3 alpha helix and has higher error on NMR observables uh than rosemary alpha. For GB3, there was one parameter that changed between rosemary alpha and rosemary alpha 1B, but this is just a minor modification of a caroxilicate bond. So, we're still trying to work out if this is a meaningful difference or more differences in sampling. In terms of next steps, we're focusing on benchmarking this force more. We will proceed with the tier three benchmarks of rosemary alpha 1B starting with protein binding for energies. We will also run an additional three replicates each of rosemary alpha and rosemary alpha 1B to uh on the GP3 target to further characterize differences between the performance on the two force fields. For the second path uh this involves a greater change to the force field. So we did some more experimentation. Our first step was to retype parameters from stage 2.1 to use the stage 2.3 partitioning although we left the vanderal parameter values untouched to retain protein behavior. We then investigated the impact of a number of variables. The first was data which data set we used for deriving initial bond and angle values via the modified simaria method as well as which data sets we use for training. Secondly, torsional flexibility. Uh whether we completely froze all torsions or refit them using prior widths of 0.1, one or 5 kilo calories per mole. Lastly, we also experimented with uh different weightings between training targets. In general, we found that the combination of variables that gave us the best performance was to use initial values calculated from larger by dance or theme data set and to train to a combination of QC archive and theme data. In addition, using the default waiting of the different training targets gave us the best results. We did observe that freezing torsions resulted in worse performance on bonds than allowing them to optimize. So we are still experimenting with allowing torsions to optimize with an arrow prior. We found that using a pri width of 1 kiloalorie per mole is enough to improve performance on bonds and tier 2 benchmarks is still ongoing to see if this will negatively impact uh folder protein performance. Lastly, Ashley is still continuing work on benchmarking the rosemary alpha force field on cyclic peptides. This quarter, the simulations have completed and she has started analyzing the results which we'll hopefully present next quarter. Um, on deliver, there have been a few updates here. As a reminder, this project is being driven by Julianne from the shirts lab and the goals are fairly similar to the protein force field. getting parameters that perform well on both small molecules and membranes and a force field that performs in line with existing membrane force fields. As a means of measuring performance, Julianne has two sets of benchmarks that she's using to compare performance. The first is alkane benchmarks uh such as torsion energetic profiles uh physical properties and transcouch ratios. The second incorporates actual membranes and measures common targets such as area per lipid blayer thickness order parameters uh form factors and she's also been looking at relaxation time. As with the protein force field, I'll just quickly summarize where we were at the start of the quarter. So on benchmarks of previous stage force fields, primarily stage 2.2, Julian found that they performed in line with other lipid force fields on metrics like area per lipid and blayer thickness. Um however they performed pretty badly on order parameter quality and were quite slow to relax. She tried a number of different refits and found that retraining to a combined small molecule and lipid data set uh didn't shift the torsional terms for alkan profiles which as you can see here get the relative energies of the trans and gash component is pretty incorrect. Retraining to an alkan only data set improved this particular target. Uh but actually matching the Q profile exactly resulted in overestimating the trans gash ratio in simulation. She also ran into repeated issues with training phosphate parameters to QM which may be because the constraint optimizations that we use uh during training with force balance result in very different confidence when optimized with mm relative to QM. The workflow that she saw the most success on was to retrain veance terms to an alkane QM data set and then retrain bands parameters to a combined alkan and methylacetate physical property data set. This improved order parameters although not to the same extent as of the lipid force fields. This quarter, following the full release of Sage 2.3, she reran her workflow and benchmarks using Sage 2.3 as a starting point. Interestingly, Sage 2.3 by itself already improves massively on Sage 2.2 with order parameters and is now performing similarly to other lipid force fields like charm. Form factor quality is still low, though suggesting possible issues with head group parameters, and the relaxation time is still pretty high. Retraining to arcane data from stage 2.3 and porting the refitted vanderal terms from the sage 2.2 fit got the stage 2.3 veance one force field. This improves relaxation time but decreases form factor quality. Interestingly this source field also decreases performance on uh general physical properties especially halogenated alkanes. Julian also tried briefing phosphate parameters using the presto library from the coal group which was a lot more successful in matching QM profiles than force balance workflows. That being said, performance on limited benchmarks definely shift substantially using these new terms um including in areas of concern. In terms of next steps, Juliana will keep benchmarking her refit force build on a wider set of liberties. To address some of the issues with overfitting to QM and physical property performance, she's planning to try copying veance of wild parameters as well as increasing the size of the training sets for her fits. Lastly, on to QM data sets and collaborations, both of which are being masterminded by Jennifer Clark. Over the past quarter, we have added a few more data sets to the QC archive. This includes a subset of molecules from SPICE 2 that we identified as problematic for various reasons, such as overlapping atoms. This is only about 2,300 compliments, but we thought it would be useful to fix for force field training and validation workflows. We also added a small data set covering PEG fragments as part of a training tutorial for fitting a force field for polymer. And lastly, the initial architecture data set for diverse organometallic lian complex structures has completed and Jen is in the process of setting up larger one for the ongoing genetic and cader lab collaboration. Speaking of collaborations, open FF has been collaborating with open FE to build a standardized benchmarks repo. I mentioned this last time as well, but just as a reminder, this repo supports relative binding for energies and absolute salvation for energies. It includes code for discovering uh available benchmark sets and includes standardized sets of partial charges to avoid charge variation affecting results. The repo also allows both building new networks and using previously generated ones and includes comprehensive samples. We're pretty excited about this project as it'll make it much easier for Open FF to run our benchmarks in the future. And indeed, Jen is actually the one who is going to be running the tier three RBFs for the protein for project. At this point, this collaboration is largely complete. There are still some odds and ends to wrap up, but most of the minimum goals have been hit. Our main focuses for the next quarter are to continue with our ongoing work, focusing on the water model refit, getting the Rosemary alpha closer to release, and continuing collaborations. We will also keep working towards stage 2.4 and helping with development of a liquid force field. So, we'll hopefully update on this project six quarter. See you next time.