Benchmarking Universal Machine Learning Force Fields with CHIPS-FF
Watch on YouTubeVideo summary
Dr. Daniel Wines from NIST introduces CHIPS-FF, a flexible benchmarking infrastructure developed under the CHIPS Act metrology program designed to evaluate universal machine learning force fields. This framework builds upon foundational resources like the Jarvis infrastructure and its extensive density functional theory (DFT) database, which are essential for training these models. While universal potentials offer near-DFT accuracy without specific training, they currently face limitations regarding interfaces, defects, and properties beyond simple energy and forces. CHIPS-FF addresses these challenges by automating the testing of various models, such as MACE, ALIGNN, ORB, and Matter, across diverse semiconductor materials to assess structural relaxations, elastic properties, phonons, thermal conductivity, and amorphous structure generation.
The demonstration highlights the framework's ability to process input files for geometry optimization, elastic strain calculations, and phonon analysis using finite difference methods on silicon supercells, producing outputs like relaxed structures, error metrics, and automated plots of bulk moduli and lattice parameters. Scaling tests reveal that ORB scales more favorably than MACE for larger systems, though current local implementations still struggle with memory consumption beyond hundreds of atoms, a limitation that GPU acceleration may help overcome. Although models like MACE are effective as DFT surrogate relaxers for sampling configurations and studying amorphous materials, their performance on complex interfaces remains limited due to data scarcity, while the tool also explores potential integration into reactive force fields to balance accuracy and scalability.
Key findings from the benchmarking indicate that while some models excel in geometry prediction, others struggle with interface work of adhesion or exhibit noise in low-force regimes affecting phonon accuracy. The framework supports a wide range of calculations, including surface energy for specific crystallographic indices, defect vacancy energies, and rapid melting simulations to generate radial distribution functions. By emphasizing community-driven open-source contributions, CHIPS-FF ensures reproducibility and continuous improvement, directly linking results to the Jarvis leaderboard to foster collaboration within the scientific community.
In conclusion, CHIPS-FF serves as a critical tool for advancing universal machine learning force fields by systematically identifying their strengths and weaknesses across different material properties and system sizes. The discussion underscores the importance of balancing computational cost with accuracy, noting that while reactive force fields like ReaxFF can scale to millions of atoms, they remain computationally expensive compared to ML-based approaches. Ultimately, the infrastructure aims to provide a robust platform for evaluating models against diverse benchmarks, ensuring that future developments in machine learning potentials can reliably handle complex phenomena such as defects and interfaces while maintaining high fidelity to quantum mechanical calculations.
Read the full video transcript
nanohub.org
online simulation and more for
nanotechnology.
Hello everybody. It's my pleasure to
introduce today's speaker Dr. Daniel
Wines. Uh Dr. Wines works as a physicist
at the National Institute of Standards
and Technology within the materials
measurement lab and his research focuses
on the integration of first principle
computational methods such as DFT and
quantum monitor Carlo um and integrating
them with machine learning techniques to
investigate quantum materials. Dr. Wines
is currently a key contributor on the
chips metrology project where he focuses
on multiscale modeling and validation of
semiconductor materials. In addition to
his research, Dr. Wines is actively
involved in the development of Jarvis
infrastructure at Nest um which provides
curated data sets and tools for
materials design. And today uh Dan is
going to be discussing one of his latest
publications um in which he will
introduce a platform uh called chipsf
which can help researchers evaluate
machine learning force fields across
materials and properties. Dr. Wines
received both his masters and PhD from
the University of Maryland Baltimore
County in 2019 and 2022 and held a
post-doal fellowship at NAS from 2022
2022 to 2024. um and he also has a
bachelor's in physics from Fordham. So
that being said, please join me in
welcoming Dr. Daniel Wines. So uh thank
you for the introduction uh Juan Carlos
and thank you for the nano hub team and
professor for uh the invitation. Very
happy to be here. Uh would love to visit
in person at some point. Uh so my name
is uh Daniel Wines and I'm a scientist
at NIST uh working as part of the Jarvis
infrastructure uh like Juan Carlos said
and today I'm going to talk about uh one
of our latest projects for uh
benchmarking these new uh universal
machine learning force fields uh with a
software package that we developed uh
called chips FF. So before I get into
things I have to legally read this
disclaimer about NIST and the chipsology
program. Apologies for how long it is
but so thank you for the opportunity to
speak or present about my research and
development effort at NIS to advocate
the micro electronics industry. The
purpose of this presentation
isformational and preliminary only. Is
not evaluative with regard to any
application or potential applications to
the CHIPS incentive program, Department
of Commerce, notices of funding
opportunities, grant challenges, or any
other department programs. Is not part
of the formal process of awarding funds.
I am not authorized to make any
commitments on any issues for the
department including on specific
questions on specific applicants.
As a funded project team from the chips
neology program, I and other NIS
researchers are presenting information
about my R&D activities with a broad set
of individuals and organizations.
Nothing that I discuss should be
understood as indicative or conclusive
regarding any shifts program. Please do
not share with me any information that
you are required to keep confidential by
law, contract, or professional
obligation. NIS, CHIPS R&D and the CHIPS
program office do not seek and will not
accept any material non-public
information outside of a due diligence
process which is separate from this
meeting. So hopefully I have more time
now to to do the rest. Uh so
like I said before we're uh our team is
working uh as part of of the chips
metrology uh program here at NIST. So
Jarvis specifically was awarded a uh
$5.2 million grant over four years uh
for a multiscale modeling and validation
of semiconductor materials and devices
project. Uh so the total chips act uh
was over $52 billion investment in 2022
signed by president Biden and our u
materials modeling infrastructure was
highlighted by the strategic
opportunities for US semiconductor uh
manufacturing a few years ago. So I just
want to introduce uh some aspects of the
project before I get into the real
technical details. So our team consists
of uh a very wide range of expertise. Uh
we have experts with density functional
theory, type binding, uh atomistic
simulations like molecular dynamics. Uh
we have TCAD device modeling, uh machine
learning techniques u to kind of connect
the dots and we have some experimental
validation as well.
So uh hopefully some people here have
heard of the Jarvis project. Uh so
Jarvis is a collection of databases,
tools. We do a lot of outreach. We have
a lot of events. So, uh I encourage
people to make an account at
jarvis.nis.gov.
So, Jarvis was started in uh January
2017 led by my colleague Kamal Chowry.
Uh since then, we have over 50 articles
published uh over 150,000 users
worldwide. Our uh density functional
theory database has over 80,000
materials in it with millions of
computed properties. So, we go sort of
beyond the standard uh information
energy band gas. we look at other
interesting properties such as uh
superconducting transition temperature
uh spin orbit coupling
uh and so on. So so kind of going past
the standard uh some of the events that
we organize are the quantum matters and
material science workshop which occurs
in the winter and we have the artificial
intelligence uh for material science
workshops. So the Ames workshop which uh
professor Strachen uh presented at last
year. So this year we're having it uh on
July 9th and 10th. So I encourage people
to register uh if you if they have the
ability. Uh so just a little quick demo
of what the the DFT database looks like.
So you go to jarvis.n.gov/jarvisdft.
we have sort of an interactive uh
periodic table view of the uh different
structures that we have in the database
where you could search by chemical
formula or ID. Um you could look at a
variety of properties in an interactive
view or you could download the data set
in JSON format to use for various uh
machine learning applications and so on.
So I just wanted to sort of get a sense
of what the interactive view of the data
is before we get into the demo of of of
using uh the data in a different format.
Uh another project uh ongoing in our
group at NIST is the atomistic lion
graph neural network uh architecture. So
basically a graph neural network based
property prediction model developed by
Kamal Chowry and Brian Nos uh which was
extended uh to be one of the first uh
unified graph neural network force
fields uh which I'll talk about uh in a
little bit as well. Uh some of the other
projects that uh we've really gotten
into in the shifts program are uh an
interface uh design toolkit called
intermat. Uh so basically this is a a
workflow to generate not just standard
bulk materials but interfaces of two
materials. Uh calculate properties such
as the band offset uh different charge
transfer and sort of analyze and
benchmark these results uh with respect
to experiment
and train uh highle machine learning
models to do these band alignment or uh
work of adhesion predictions as well.
Uh one other project uh we have a tight
binding code um as part of our chips
program uh chips project. It's a
predictive type binding model. Uh some
of the features include uh charge
self-consistency two and three body
terms in the Hamiltonian. Uh and what
makes this uh type binding code unique
is that there's full coverage of the
entire periodic table uh trained on
massive or fitted to massive DFT data
sets of millions of structures. Uh so
there's been extensive uh DFT database
fitting and testing sort of in an active
loop uh and we get very nice performance
for for this as well. So a useful tool
for extending beyond the scale of DFT.
So why am I talking about all these
things? So the main motivation uh of our
project is multiscale modeling and you
can see the classic multiscale modeling
diagram where we're going from the
quantum scale uh using methods such as
entity functional theory all the way up
to the macro scale and what uh in this
talk what I'm particularly focused on is
bridging the gap between the quantum
scale accuracy and atomistic simulations
and uh of course one way to do this is
through machine learning methods. So the
concept of a universal pre-trained
machine learning force field is
relatively new. It was uh started in the
group of Shu Pingong uh a couple years
ago uh for the graph neural network
based machine learning uh pre-trained
force fields and there's a lot of
different names for them. There's a
machine learning universal pre-trained
un uh machine learning force field
machine learning interatomic potential
uh people are calling them foundation
potentials now. So they all mean the
same thing. So basically uh graph neural
network architectures that are trained
on massive DFT databases and uh geometry
optimization trajectories
uh which are hopefully transferable to
to a wide variety of systems and have
comparable accuracy to DFT. So one of
the main advantages is uh they utilize
large DFT databases that are existing.
So such as data sets like the materials
project or uh Jarvis DFT that encompass
uh all the elements of the TBR table uh
so they're transferable
uh you can use them sort of out of the
box kind of like a black box so no
training or DFT generation is required
or training and you could achieve uh
near DFT accuracy for certain properties
especially for uh structural relaxations
of bulk materials. for energy and
forces. So some of the disadvantages are
most of these if not all are are trained
on uh ground state bulk structures at
zero kelvin. uh there's no defects or
interfaces
and also uh the predictions can be much
more expensive than classical force
fields uh such as something like a
reactive force field or or something
even more simple like a Leonard Jones
force field which can scale to maybe
hundreds of thousands of atoms pretty
well. uh and also uh scratch training
and fine-tuning these uh universal
foundation potentials can be expensive
and difficult and the performance for
properties uh beyond energy and force
are still remaining unclear
uh which is sort of the point of this
work. So within the past two years uh
several of the universal machine
learning force field frameworks have
been developed. There's a strong
interest from from industry on this as
well. So almost every major company is
working on this now. So I'm not going to
get super into the different
architecture choices. There's certain
choices that are made with how these
models are trained uh versus training on
forces directly or computing forces uh
using invariant versus equivariant
uh architectures and the message
passing. And it starts to get very
technical here. But just the point is
there are several different models all
of of varying flavors. Uh some of these
are proprietary.
So uh companies such as Microsoft,
Google,
uh Meta, uh some startup companies like
orbital materials or preferred networks
all kind of have a stake in in the game
of these potentials now. and the the
degree of how uh they're handling the
open access versus closed act access
depends on the company. Some are all
open access, some have proprietary ones
beyond payw walls. Uh so for our
interest we're interested in
benchmarking the open access potentials
and which leads me to my next point of
sort of these DFT data sets for uh the
machine learning force field training
and for those of you that don't don't
know it's been a extremely huge year for
this field just in terms of basically
since 2023
none of but only one or two of these
existed and now we have dozens and in
terms of data sets there's been uh a lot
of of uh leaps in releasing massive DFT
data sets of of millions of structures.
So before we had uh sort of working on
the level of tens or hundreds of
thousands with uh databases such as
Jarvis or the materials project. There's
also OQMD uh other data sets as as well.
uh but recently we have the Alexandria
data set uh which has been released
which has uh very high quality data of a
variety of materials on the order of
millions of structures and we have the
open materials uh 2024 OMAT 24 data set
from meta uh which has hundreds of
millions of uh different snapshots of of
trajectories of molecular dynamics or uh
density functional theory relaxation. So
these all been very very useful for for
the community. I think even since I've
made these slides, there's been uh
another data set related to chemistry
released by by meta by by the same team.
So the OML 25 data set I think it's
called. So this is very rapidly evolving
things are changing by the week. Uh so
what I said in addition to these open
access data sets we have uh some other
data sets like the uh proprietary ones
uh from Google. So the genome uh data
set and model which was published in
nature about two years ago. Uh we have
the matter DFT data set and also a data
set from the Japanese company uh
preferred networks.
And very recently from uh Berkeley lab,
we have a uh very high quality data set
that's been released. Uh so basically
what they call a foundational potential
energy surface data set for materials.
Uh encompasses about 400,000 uh
snapshots of very high quality data
which will hopefully be used in uh
future machine learning force field
training. So with the with the onset of
all these different data sets, all these
different models, uh benchmarking is
becoming very important and uh one of
the first efforts to actually capture
how these potentials are are doing on a
targeted task uh was the matbench
discovery framework uh which was started
about a little less than two years ago.
Uh so it's a very popular benchmarking
platform uh especially for the universal
machine learning force fields. Uh some
of the focus is on uh so how energy is
captured uh things like formation
energy. Uh they've added thermal
conductivity predictions recently and
they have very large scale test sets uh
usually a subset of the materials
project of about 200,000 or so
computations that need to be done for a
user to submit their benchmark to here.
And this was a snapshot from a few
months ago. I'm sure this is already
different. it it's changing sort of by
the week or month. Uh so in addition to
the mapbench discovery u uh at NIS we've
developed the Jarvis leaderboard uh
which is a similar goal. So we have a
large scale benchmarking effort to sort
of address uh some of the other issues
in the material science community. So
we're very concerned with
reproducibility
uh transparency of data validation the
fidelity of the data
sort of uh what the data versus the
metadata might be in a simulation uh
sort of how we can compare what we get
in a model to the ground truth and we
really want synergy of uh the
computational and experimental
databases. Uh so we published this paper
last year in MPJ computational
materials. Uh so the driver's
leaderboard large scale benchmark of
materials design methods and we have an
interactive website uh so a NS website
which hosts these benchmarks uh that are
very dynamic. So some of the benchmarks
that we host on this website and you can
see a list of all of our uh
collaborators here. If Arun is in the
audience he he helped collaborate on
this project as well. Uh so we're we
we're testing different uh methodologies
such as electronic structure methods uh
so that's density functional theory
quantum Monte Carlo GW
various artificial intelligence uh
techniques so graph neural network based
single property prediction
multi-property prediction image
classification generative models we're
looking at classical force fields so uh
classical force fields and machine
learning force fields and we have uh uh
a small subset devoted to quantum
computation algorithms and experimental
data as well.
[Music]
So this is an open platform for for
people to uh submit their benchmarks to.
So for example, if there's a certain
target uh property from DFT and we have
we train an AI model to uh reproduce
that DFT, we can sort of submit
everything we want to submit uh a script
to completely rerun our model and test
it metadata related to the software
versions, the hardware versions, um any
relevant information to to reproduce
this. And then once it's uploaded to the
leaderboard, um errorometrics can be uh
recorded and then basically tracked over
time. So here is just an example of a AI
formation energy per atom uh with a a
variety of different uh machine learning
methods. Some of them are
descriptor-based, some of them are graph
neural network based. And thing is a as
time goes on and and these uh models get
better or they're trained on better data
uh it it's important to keep track of
this and we sort of have a a a record of
how we can move forward and push the
field uh in this direction.
So, here's just another kind of like a
zoomed in version, a snapshot, and so
far we have over uh 300 benchmarks, uh
about 2,000 contributions, uh about 30
contributors,
millions of data points, and uh we're
starting to encourage people that when
they publish a paper that you don't just
have, uh data availability or code
availability, but maybe actually have
something uh concrete like a benchmark
availability uh that can be associated
with your manuscript and tracked over
time as as different models are are
developed. So I want to shift gears and
and go back to the uh the machine
learning uh intercomic potential. So the
universal potentials and uh talk about
some of the more
uh recent and focused benchmarking
efforts. So with the rapid development
of these potentials uh it becomes even
more
uh relevant to sort of go through the
process test these potentials for
certain applications uh see how they do
if they could even go beyond bulk
material uh sort of structural
relaxations.
So there have been some uh benchmarking
studies related to phonon surfaces
thermal conductivity.
Uh so as this is changing by the day
there there's more uh in-depth ones like
there people have looked at uh combining
universal potentials with calad base
phase diagrams for for alloys
uh looking at how force constants are
changing with different potentials and
so on.
So why is benchmarking of universal
potentials needed? Uh this can really
allow us to make a more informed
decision on which potential to use not
just in terms of the of the accuracy but
of the cost as well. And
for for people that are familiar with
DFT uh so there's sort of the the
Jacob's ladder picture of DFT where we
have different rungs of the of the
ladder. We start with archery faul which
has uh no electron correlation and we go
to chemical accuracy which is sort of
the heaven of of how accurate a DFT
calculation can be and as we go up the
rungs of the ladder um we become more
accurate but of course the cost goes up
and there are certain systematic trends
in in this ladder that we've discovered
over time such as LDA versus GGA which
could be underbinding or overbinding or
or predicting lice costs that are too
big versus too small. So we we sort of
want to try to understand this for
universal machine learning potentials
and it can help us understand the
applicability what accuracy we might
need for certain applications and help
us to develop uh better force fields in
the future.
So to do this uh we developed a
infrastructure called the computational
high performance infrastructure for
predictive simulation based force
fields. So chips FF which is a a nod to
our uh chips funding and chips metrology
project and our goal is let's create a
userfriendly very flexible benchmarking
package to test various machine learning
force fields for any material. Uh so we
added some capabilities that go beyond
sort of what we were seeing u like
structural relaxation. So we want to go
beyond structural arrestation look at
elastic properties defects uh surfaces
material interfaces
uh phonons thermal conductivity
molecular dynamics and have a very easy
pipeline to automate uh the error
metrics as we do this.
So here's sort of a workflow and we we
recently published this paper uh in ACS
materials letters uh earlier in 2025
myself and Kamal Chowry also at Nest. Uh
so sort of the workflow how this works
is we have universal uh potentials which
we connect with uh the atomic simulation
environment so ASSE and our Jarvis tool
software and we have a variety of of
different uh properties we can look at.
So we can look at optimized structure,
bulk modulus, elastic properties,
phonons, thermal expansion, thermal
conductivity,
molecular dynamics, um so we can melt or
quench structures, look at amorphous
materials, look at the vacancy energy,
surface energy and also look at
interfaces. So in our case uh we can we
can take starting structures from Jarvis
DFT to feed through our pipeline uh to
to manipulate them or or create
different vacancies or surfaces or
interfaces or or even test on both
materials and then we can use the data
that we have in our database uh to
compute the error metrics. So, MAE or R
squar with respect to uh what we're
testing on for these potentials and then
create uh create them into a format
that's appropriate for the Jarvis
leaderboard and automatically uploaded.
So, here's a a little bit of a snapshot
of our repository and we we'll look more
closely uh in the NanoHub tool of this
uh that
uh myself, Juan Carlos, and Rook Shik
helped uh to get online a couple days
ago for this workshop. Uh so if you go
to uh github.comusnisggovchipsf
uh you can see more detail details about
this. I encourage people to go check it
out, give it a star or a fork if they
want. Uh so basically the way this works
is we we wanted to sort of model this
off of using a a regular atomistic uh
software like VAS for lens where we have
an input file with some designated tasks
parameters to tune. Um, we have an
easytouse command line interface. Uh, we
could easily parallelize this for HPC,
whether that's on CPU or GPU,
where maybe instead of running one or
two examples, we want to submit a
calculation for for 10,000 materials or
100 different force fields. It's all
very doable. We have a couple
interactive notebooks on the git github
page uh which are similar to what I'll
show in in the nano hub tutorial uh at
the end of this. But here's sort of like
a snapshot of what a full input file
might look where we're giving a material
ID. Uh we're giving the type of
calculator. So in this case the charge
net universal potential uh where certain
files we might need like chemical
potentials uh different properties to to
compute uh different relaxation settings
or uh supercell sizes uh displacements
and so on.
So here's sort of a snapshot of how uh
some of the some of the functions the
main functions that are used and on the
left is sort of how we're loading in uh
the different calculators through ASC.
Uh so the the way that this is set up
makes things very easy to add a new
calculator over time. And basically it
really really all it takes is three or
four lines of code. So for example, if a
new force field is developed tomorrow, I
can add three or four lines of code uh
update one file and then test all the
materials I want for a new force field.
So we made this pretty robust uh in
terms of that. Same thing with version
upgrades. For example, if there's a bug
in an existing force shield and they
have a new version, we could easily
update that run through the same
pipeline and same thing. So here are the
main functions. Like I was saying, we
have relaxation for optimizing atomic
structures, uh formation energy
calculation, elast elastic tensors,
uh calculating energy versus volume
curves, uh phonons, defects, uh phony
for third order force constants, uh MD
or intermat simulations for interfaces
as well.
Uh another thing we were interested in
is looking at the scaling o of these
different models and uh this can can be
pretty important in terms of deciding
which model to use uh outweighing
accuracy versus uh scaling. So we we did
we actually did the scaling for CPU just
because we were interested in seeing how
a lot of people are don't have access to
GPUs especially in small environments
like looking at Google Collab or or or
on your own desktop maybe. So we want to
see how how do these force fields
actually scale the CPU and we see uh
good pretty good scaling with some of
the the older force fields that are
maybe not as accurate uh such as the M3G
net MattGl force field. Um some of the
equarian force fields like mace or 7net
have a have a a little bit higher
scaling. uh the
OMAT models which I talked about so the
the models that were recently released
from Meta have pretty unfavorable
scaling with the system size but as
you'll see in the next few slides they
are the most accurate for a lot of
properties.
So in our case instead of benchmarking
on a extremely large uh data set like
something from the materials project
that map pench does we wanted to focus
on a small subset of materials and look
at more properties rather than just a
larger set and look at just force or
energy. So we looked at a pretty diverse
data set of around 104 materials that
are common for semiconductor devices. Uh
they have a pretty uh diverse group of
formulas. There's some metal, some band
gaps, some insulators,
uh
pretty diverse crystal system and we
have some bulk and some sort of quasi 2D
verander wall materials as well.
So we went on to to look at uh a variety
of different properties like I said and
we benchmark with uh respect to the
Jarvis DFT database. So the Jarvis DFT
database uh all the the data is
calculated with the OP B88 Vanderwals
functional instead of the standard uh
GGA or PBE which is used in a lot of
other databases uh without a Vanderwal
correction
and and we're we're seeing that uh
consistently the the OMAP models and the
ORB models from orbital materials uh
with and without a Vanderwalls
correction are are producing very
accurate geometries.
Uh a line force field is doing a good
job as well, but it's possibly biased
that we're starting with the same
structure that we have from Jarvis which
align was trained on from the so it's
trained on the vendor walls material and
is predicting vendor walls corrected
lattice constants. uh we have less luck
with some of the elastic properties and
bulk moduli uh which is in an indication
that that some of of the out of
equilibrium uh properties or potential
energy landscape is not as accurate for
some of the models.
So we also looked at uh more complex
properties. So we looked at surface
energy and vacancy energy. And basically
the way we we do this is we construct
different surfaces with uh various
terminations and we take an energy
difference with respect to the bulk and
divide by the cross-sectional area. Uh
and we're comparing also uh data that
has been calculated in Jarvis. So we
have a large vacancy data set and we
have a large surface data set as well uh
for a ground truth comparison of our uh
machine learning force field asse
workflow.
And uh similarly for vacancies we we can
look at the difference of the energy of
the defect cell. So create a large
supercell remove an atom with respect to
the bulk and chemical potential. Uh so
we're looking at sort of neutral
defects. So no changes of charge states
but currently that's what we're limited
by uh for the machine learning force
field frameworks and we see pretty
consistently that uh orb and the the
OMAP models are doing a very good job
with services and defects. Some of the
other equavarian models uh such as mace
or or uh 7net also do a good job as well
as the Microsoft
uh Madison force field also does a
pretty excellent job with with uh
vacancy formation energy.
So in addition to the open source force
fields uh recently
uh there's been a force field released
by preferred networks. So what's called
the PFP potential they have uh basically
taken our code and ran the workflow with
their potential. So it's proprietary
it's not open source and they were able
to get uh very comparable results to
some of the best predictions that we've
had for surface energy which was very
nice finding from them that they
published in their blog about this. Uh
so very interesting
some of uh the limitations of the
machine learning universal potentials
have been for interfaces.
Uh so what you see here is we're looking
sort of at two different materials
something like aluminum 111 aluminum 203
001. And we want to calculate the work
of adhesion between the two materials.
And in addition to this, we want to sort
of do an XY plane scan of what the
energy landscape looks like uh for minor
pertabbations of of of scanning the two
materials together. And the XY scan can
give us a little bit of an idea of how
smooth uh the potential energy surface
might be at the interface.
So what we found is that and and like I
said before uh we were using these
machine learning force in conjunction
with uh the intermat package for
interface design and generation which I
talked about at the beginning. So we
have a pretty seamless interface uh no
pun intended between the the uh
intermath software and chips FF uh for
for calculating these properties as
well. And we find that the work of
adhesion predictions are pretty poor for
the universal machine learning force
fields. Uh when we're doing XY or Z
scans, uh there's some variability in in
what the potential energy surfaces look
like or the energy minimums look like.
And I think this is a room for these
models to grow. Uh so like I was saying
before, a lot have not been trained on
interface data. And I think that in the
future this might be a good opportunity
to sort of fine-tune or or or make the
models better.
Uh one other interesting finding we
found was uh for phonons. So we
calculated uh the the properties of
phonons for about uh 85 of the 100
materials
and sort of seeing how how good can we
get the force const constants and sort
of the overall error in the funon band
structure uh using the finite
displacement method.
And what we see is that interestingly
uh the displacement that you use in the
finite displacement method I is very
important for the accuracy that you get
uh in terms of this. So ideally you want
to use as small of a uh displacement as
possible. You really want to capture the
anharmonic phonon modes. Um, but we
actually see that when we use high
displacements, which we would think
would be not realistic or inaccurate,
they're giving us sort of a false sense
that these models are accurate,
especially for the OMAD models and the
ORB version two models. And as we
decrease the displacement to what really
it should be, uh, the error drastically
increases. And this signifies that there
is uh some significant noise in the low
force regime and it possibly might be
due to how forces are computed. So for
these models uh the forces are computed
by are by learning them directly and not
computed from the energy surface. So
it's through something like back
propagation.
So this has been a discussion in the
community uh especially for the folks at
or renewable materials that develop this
model and they've actually released a
version three of Orb which uh I think
sort of solves some of this issue. Uh we
have not tested it through our pipeline
yet but we plan to do that very soon. Uh
so there is a new version of orb that
has been released since then.
Uh so we have more consistent phonon
predictions uh through different
displacements uh using some of the other
state-of-the-art potentials. Uh very
good results for the matter from
Microsoft.
Uh one one interesting result that we
also tested was can these machine
learning force fields uh universal
potentials actually be used for uh
creating amorphous structures. So
basically melting the structure and then
quenching it at room temperature. Can we
actually recover the radial distribution
function of an amorphous silicon? And we
find that they all do a pretty
consistent job of doing this. uh we
found the best results for across the
board for the many other properties like
we said for the all models orb and
matter and we're comparing uh directly
with abonio molecular dynamics which is
orders of magnitude more uh expensive to
do this so I think this is a very good
application of using uh universal
machine learning potentials for actually
getting a starting structure for
something like an amorphous material it
really can cut the computational effort
to get towards that goal. And this is
some we're doing some followup work
right now sort of benchmarking this for
more and more amorphous materials uh
looking at other properties such as uh
the density and so on. So keep an eye
out for that. We'll we'll probably be
working more in that space very soon.
So I just want to go over the last
little bit of this quickly before we get
to the tutorial. Uh so like I said we
were focusing on a smaller data set but
we wanted to test a couple properties uh
automatically with larger data sets. So
we looked at the ML data set which
consists of a couple hundred uh binary
crystals uh looking at the force error.
So we have the capability in chipsf. So
if you want to you can be flexible in
how what data sets you want to test uh
whether it's a handful of materials or
something like on this slide where we
have force error for uh
data sets that were used to train force
fields such as the one for align or the
one for magnet or uh materials project
which was used to train mace or uh seven
and many others.
So uh finally really the main takeaway
here is we have a very dynamic way to to
keep up-to-date benchmarks now hosted on
the Jarvis leaderboard. uh we have sort
of a special benchmark here for chips
fep which consists of all the different
models and the different properties of
here's like the pearing coh coefficient
here and we can look at the more larger
data sets as well like mlearn to see the
force error of of different models and
we plan to continuously change this over
time as new models are being developed
already on my radar is to add orb
version 3 and some of the others I think
there was another model that was
recently released by the startup company
uh radical AI and other various startups
are are taking their stab at making very
accurate universal potentials that scale
very well. So I just want to uh
emphasize that community contributions
are very important to projects like this
especially at NIS we're very uh focused
on open source and obviously that's the
main mission of NanoHub as well. Uh so
we've had some really nice contributions
actually from some of the developers of
these force fields that are working in
these companies. Uh so for orbital
materials for for example uh we were
able to find a bug in one of the
releases of orb due to this when uh we
looked at our amorphous structures and
then they were able to make the
appropriate changes we updated it we
reran the benchmark and we got very nice
results. So that's the beauty of
community contributions archive GitHub
and and all of these great things. So I
just want to end with uh we we
introduced chips FF our benchmarking
workflow for universal force fields uh
benchmarked on a small subset of
materials for for applications in the
semiconductor industry uh linking this
directly to the Jarvis leaderboard. We
plan to extend this for for some more
complicated properties uh like thermal
expansion. We want to extend this to
alloys
uh of course benchmark this more for
morpher structures. And another thing
we're very interested in is uh
benchmarking dislocations as well of
materials. So now that some of these
force fields have more or less okay
scalability, we think that's also a
realistic goal for for what we can do.
And with that, I think I could take a
couple questions before we jump into the
tutorial. Uh, and I'd like to thank
everybody for their attention. Thank
you. Thank you, Daniel. Uh, this is a
great presentation. All right. So, we
have a couple of questions here. Um,
first question here was about your um
party plots. So, the question is, most
of these models are trained against PBE,
but here you're comparing with different
functionals and calculating error
metrics.
If the model can only be good as good as
the underlying functionals, what um do
the error metrics are trying to tell the
audience?
Yeah, I think that's something you have
to be careful about and it's it's been
sort of a debate for a lot of these
things. So, the flavor of DFT for all of
these databases is different and uh I
think our main goal is not to say like
okay this is the end all be all metric
of something. I think it's more to say
we should be flexible in how we
benchmark things and depending on the
property that might change but uh we we
have the capability where you could
change the underlying uh ground truth to
something like Alexandria uh materials
project even experiment if you have the
experimental results as well. And from
from working in the space for a little
while, I I honestly think it is maybe
not the most productive to chase very
small increment differences in MAE. For
example, even for most of these
properties, Vanderwal's DFT and PB will
give pretty close results depending on
on what it is. And I think that most of
the time the error of the force field
will be within the regime of what the
error between those two types of
functionals might be. So I I I think
the short answer is we should be careful
with how we assess how good or bad a
force field is and make informed
decisions based on that. see depending
on what this was trained on, what this
was compared to and maybe what is
actually reality.
Okay, that makes sense to me. We have a
question that is kind of building on
that. So when you were mentioning your
errors, you were saying that these
errors are good or good enough. Can you
comment on uh what you meant by good
enough? When is an error good enough for
your parody plots?
Uh yeah that's a good question and I
guess it's very arbitrary but I mean I I
think at the end of the day we still
cannot rely on these models to be the
end all be all situ uh solution of of a
problem. I think they could aid in
screening for example like if you are if
you want to calculate a defect energy
for a particular system maybe you relax
with the machine learning force field it
gets you closer to the answer but then
you finish it with DFT and I think when
we're getting into
small energy differences like for
example for surfaces a surface energy is
on the order of let's say like two jewel
per uh meter squared But we're finding
error that's about two orders of
magnitude lower than that. And same
thing the defect as well. I I think that
that's
the best the state-of-the-arts can do
that are trained not on surfaces or
defects. Of course, we we have I think
uh
I guess stricter error uh comparisons
for things like lattice constants and
and formation energies. And it gets to a
point where the NAEs that people are
showing like okay this we have an MAE of
0.00
one or something in the last concept 05.
At that point I think we're almost in
the error of what DFD can even get you.
And I think at that point it's better to
kind of just say okay all of these
models are doing a very good job. which
is why I kind of was careful not to just
highlight one model that did the best,
but sort of highlight multiple green
rows of things that are doing very good
for a particular test set like that. So,
I I I think same thing with the
amorphous structures as well. I think
once we're we're getting into sort of
these very minor differences, I think
it's safe to say that they're all doing
a good job.
Okay, that makes sense. We have quite a
few more questions, but I think it would
be good if we could run through the
tutorials for a few minutes and then we
can go back to the questions. Yeah.
Yeah, sounds good. Yeah. So, I don't
have a super long tutorial, uh, but I
just kind of wanted to give people a
sense of of uh how how the software
works and uh feel free to visit the
GitHub page, uh, clone it yourself, uh,
play around with it. if there's new
models you want to add and so on. So
yeah, I want to thank uh Rushik and Juan
Carlos again for helping you get this
set up. So I think uh I I kind of just
have an example of of that workflow I
was showing where uh we're going to run
a variety of different properties uh
look at some error metrics with respect
to to some of the ground truth data that
we have imported. Uh so in our case
we're not going to download the whole
Jarvis DFT database or materials
project. I just provided uh some JSON
files that have the structures that
we're interested in to just sort of read
in get a starting structure and then
compare the properties. But this is a
very flexible arbitrary thing. You could
change this comparison or starting point
to to whatever you might want. So we
could start by uh sort of importing the
necessary functions we have and then we
could uh read in the JSON files of the
uh sort of the demo DFT uh bulk
materials vacancy calculations and
surfaces.
So in our case this test set consisted
of over 100 different materials for uh
semiconductor devices. So that's what's
in these as well. and we can go ahead
and we'll run a a sample geometry
optimization. Uh so so we have a couple
different uh I guess environments that
can be run here on NanoHub. So we have
the chips FF1 and there's a chips FF2.
Uh so one thing to to be careful of is a
lot of these machine learning force
field uh they do not work well together
in the same environment. So, uh, I know
Juan Carlos and, uh, Roshik
painstakingly tried to get as many to
work in each one as possible, and I I
give them a lot of credit for that
because I I'm the person that will just
make a different kind of environment for
each one. But, so I appreciate that a
lot. So, I think this one has
orb, mace, and align, and then the other
one has some of the other models like
uh, Matter Sim, and I think Chargenet.
Correct me if I'm wrong. Uh but so we we
can look at uh a series of different
input files uh for four different
property predictions that we have. Uh
all of these input files are available
on the git on the GitHub repository as
well. If you go to chips FF and example
inputs, you'll see them all here and
they can all be modified to what you
want. So first let's run a simple
geometry optimization using orb version
two and mace and then apply some
isotropic strain. So energy versus
volume and we could do this for
something like silicon and gallium
arsenide. So this is what the input file
looks like. Uh we have the different
Jarvis ids that we want to run. We have
the different calculator types
properties that we want. So we want uh
to relax. We wanted to calculate the EV
curve uh some of the relaxation
settings. And then we can go ahead and
run this right here. Hopefully this will
take just a sec. So it's loading the
local data set and then it'll go ahead
and uh optimize the geometry of both
materials.
So see starting the relaxation with orb
version two.
This should just take a sec or two.
Yeah. While this runs, we'll we'll just
go through the rest and I guess
afterwards people can can play around
with it as well. Uh yeah. So after this
runs uh there there's sort of a series
of of uh output files that can be
generated. So here we have PNG file of
the E versus Vcurve also in text format.
We have the relaxed structure and POSCAR
form format and some of the error
metrics as well uh in JSON and CSV
format. So we could plot with orb what
the E versus V looks like and the bulk
modulus of gallium marsenide. It does a
very decent job of of calculating bulk
modules here. Uh and we can go ahead and
and calculate the uh the mean absolute
error of uh what these might be. So once
that finishes, you'll be able to sort of
make a little bit of a parody plot of A,
B, and C for the DFT versus uh the
machine learning force field calculated
lattice parameters and it will
automatically calculate the uh constant
the the errors and lattice constant uh
for orb and mace with scikitlearn.
So I don't think my plots will appear
unless my my calculation finishes. But
you can see here we have lattice error
for orb of 012 angstrom for AB and C for
those two compounds and for mace we have
a similar error of 015 and of course you
could extend this to to a larger uh
scale computation if you want. You can
you can do this for a number of
materials uh or number a number of
different calculators as well.
So what else? Uh we have uh the input
file for for a phonon calculation. So
this is example calculation down here.
While the other stuff's running I can go
over uh so we have example phonon
calculation with or version two uh using
phonop using the finite difference
method to get the force constants for
silicon.
And if you see here what we need to do
is we need to relax the structure of
course first and then run the phone on
analysis. Uh so we're doing a 2x twox
two supercell uh and we can pick our
displacement. I had a discussion about
the displacement before of of how this
can change depending on the force field
and how forces are are computed. And
after this runs, it should only take a
couple minutes. So basically we're
calculating the force constants for the
different supercells. Uh we have outputs
of uh force constant files that you
would normally have in phon which you
can do whatever you want with. uh YAML
files of the different band values that
can also be used for for uh I guess if
you wanted to compare the error in the
highest frequency or the lowest
frequency uh that that's what a common
thing that people have done and and we
can go ahead and look at the uh phonon
band structure as well and density of
states from this uh in a very automated
and easy way. So it's it's generating
these plots automatically when you run
them. Uh we could also look at some of
the other thermal properties uh such as
the uh
zero point energy which could be
important in in determining uh the
stability of a compound. We could look
at the heat capacity and the uh entropy
as a function of temperature as well.
Uh yeah we have a couple we have a
couple other examples here uh for doing
surface and uh defect calculations. So
for for surfaces what what we can do is
in the input file we can define uh sort
of what what index for the for the
substrate I what index for the surface
we want whether it's 001 or 1 0 0 uh and
then we have an automated way to
calculate the surface energy so in the
input file we have uh computations of
the uh bulk energy and also the surface
energy. Uh so there there's different
uh you can adjust the vacuum, you can
adjust uh the indices you want here. So
if you see here the number of layers, if
you go to the surface settings, uh we
can say do we want the volume to be
constant or not. uh the the different
indices like I said for 001 111 so on uh
how much vacuuming we want to give 18
angstrom 20 angstrom so on so uh we
could run this and we could do the same
thing and once this this uh finishes on
your guys end you you can plot the
parody plots of uh for orb here so I'm
picking orb uh just because orb is uh
one of the more favorably scaling force
fields that that we could test things
on. Uh, of course you could pick
whatever for force field you want.
So we could again we calculate the MAE
of the surface here uh using orb. So we
have MAE of 295 jewel per meter squared
for for 001 silicon. Similarly we could
do this for uh defects in silicon as
well and and get our par plot and
vacancy energy. Uh one other simulation
we can run is the uh melting and
quenching simulations for amorphous
silicon. So uh what we can do here in
the in the input file is we can define
things uh such as of course the which ID
we want to run it for our time step how
big of a supercell we want. So we have
an eight angstrom supercell here the
number of steps for the melting number
of steps for the quenching. Um, so we
want 6,000 Kelvin, 300 Kelvin, and then
we can sort of see in ASSE how this is
done where we're we're melting for a
certain amount of time. It's a very
these are not production level
calculations. This is very little quick
tests, but in the end, we'll output uh
some things such as the the quench
structure and posar format format and
the radial distribution function, which
you could visualize. And of course, this
is not amorphous. There's a lot of sharp
peaks here because we only ran the
simulation for a couple time step test
sets. But if you if you run the
simulation out long enough, you will get
what you're supposed to get. And the
last little thing I want to uh show you
is we have uh some examples of testing
the automated scaling versus the wall
time of a copper supercell with various
machine learning force fields. And I
think this is very important in terms of
like I said testing which force fields
you might want to use outweighing
accuracy versus scalability. So we have
like a 1x one by 1 3x3x3 6x 6x6 so on
supercelly here uh using orb and mace
and we can get a sense of how these
scale uh looking here. So, uh, orb
scales much more favorably favorably
than mace here. And,
uh, you might have a more consistent
trend the bigger you go. If you go out
to a couple thousand or 10,000 like I
was showing in in our paper, but you can
kind of get a sense of if you're work
whatever system range you want to work
in, you can see how each force field
will will change.
So, uh, like I said, there there are a
variety of different force fields in
either of those notebooks, uh, for
so I think, yeah, the variety of
different notebooks I encourage you guys
to check out, uh, maybe a little bit
after this once some of the traffic uh,
calms down a little bit. So we have orb
mace and align FF in chips FF and then
in chips FF2 we have Madison and
ChargeNet and if we have time for maybe
a couple questions I' happy to take
more.
Yeah we have uh we have time for a
couple more questions. There's there's
quite a few here but let me pick a
couple of them. Okay. So yes so there's
a couple of questions that are asking
about system sizes. what are the system
sizes that these potentials can handle
efficiently? Like what would be the
largest simulation that you can um
foresee with these potentials? Yeah, I I
think that's a huge bottleneck of what's
going on right now. So
one thing is the scaling, but another
thing is the memory consumption, which
is what you're seeing right now.
Basically, these consume a lot of memory
and even on if you're running on CPU and
you have a pretty sizable node with a
lot of memory on it, these can still
crash or run very slowly. Uh, for system
sizes of maybe over a thousand. I think
in the hundreds range, it can handle
pretty good. But there have been a lot
of very recent efforts I'm seeing from
uh at least I think the sketer group in
in Berkeley and some uh some of the
startup companies that I was talking
about such as this radical AI company is
they're getting these force fields to
run much more effectively on GPUs now
where I think it's more feasible to run
more of a molecular dynamics next level
calculation of maybe a couple thousand
atoms where I think if if you're just
running these locally out of the box how
they are here maybe you could handle a
little bit more than DFT. It's feeling
like it's maybe a DFT or abonio
surrogate as of now. I I think we're not
quite there with the scaling these up
yet. So hopefully with some more recent
implementations and uh some of the the
new codes I guess people are switching
things from ASC to PyTorch now and and
and such. So I think we'll hopefully be
able to scale up and use the same models
in a different framework like that.
Okay, that makes sense. There's a couple
of questions too on OMAT. So um one of
the audience members is asking
um as far as they understand OMAD was a
maze model. Is there a difference
between the OMAD model and the other
maze models? Is it different training
set or architecture? So uh so there
there's a little they're a little bit
different. So
uh
they're both equivariant neur neural
network potentials. Mace is an
architecture that's based off of atomic
cluster expansion and there's different
modifications that were made to it after
that. I think yeah OMAT is a
sort of a spin-off of this equivariant
dense former model or something like
that. So they're they have a similar
architecture but they're not exactly the
same.
All right. And also in OMAD um another
one of the attendees is asking for your
views on how OMAD predictions are on
interface thermal properties in 2D uh
materials with defects.
Yeah I think have any comments? Yeah I I
think that we're not quite there with
doing super complicated things with
these potentials yet. like there they're
there's a lot of work to be done on the
data side and probably on the the
finetuning side. Uh interfaces are are
not super great but I think the best
thing that these potentials could be
used for are as a surrogate DFT relaxer
for getting accurate energies and
geometries.
So we
and that's not saying that we might not
get there eventually, but I think like
for example, if you want to sample a lot
of different configurations and get a
lot of energies, a lot of structures,
this is a good method. It's showing
promise for the amorphous materials, uh,
which hopefully we'll have more about
that soon. Interface wise, I don't know
if we're quite there. I think that there
needs to be a larger interfa larger and
accurate interface database that
hopefully can be used to train uh newer
models or fine-tune them which would be
good but of course getting an interface
data set is tricky. I mean it's very
costly with DFT and there's a lot of
assumptions and caveats with that as
well. But I think
hopefully we will get there. I mean with
the pace that things are moving, I
wouldn't be surprised if within a year
or so someone did that.
All right. And then I think we have two
more questions and then we can we can be
on our way. So one of the attendees is
also asking if you know about any
ongoing work towards making these types
of force fields but reactive like reacts
FF
like universal reacts FF potentials
basically
I am not aware of any of work like that
but I think it would be interesting to
sort of incorporate machine learning
into the fitting of reacts potentials. I
know some of my colleagues, we've
discussed this before a little bit, but
yeah, that's maybe a little bit out of
my domain with React potentials, but I
think that I guess like React is is
reasonably accurate and it can scale
much better than this. So, I think it
would be an interesting route for for
someone to go maybe. So, so just a
comment uh on our group, we developed a
a potential a reactive potential for
CHNO
uh based on machine learning that
actually outperforms the the React FFS
that that we've been uh using for for
two decades. So, so there's it's
challenging but but there's some some
work and it'd be nice to uh extend the
work you guys are doing to uh reactive
intraatomic potentials for molecular
systems. Yeah. Yeah. No, that's that
that's great. And I I think that's
definitely
an important push like in at least
scaling up what you can do cuz for net
like obviously there there's you can't
really go what's the size limit you can
go with reacts uh professor is it like
in the 10,000s or can you you can go up
to 100,000?
Yeah. No, React scales very well. uh
we've um I don't know the biggest that
we've done but uh uh more related to
access to large computers than anything
else. It it scales linearly with system
size. Uh so we've we've run maybe 100
million or 10 million atom simulations
with react FF and um
it's more expensive than simpler
potentials but it scales very well and
it is implemented in Kauos for uh GPUs
so it you you can actually run big stuff
with reacts. Yeah that's great then yeah
uh with that thank you to uh Dr. Wines
for the good presentation and thank you
everybody for attending.