Deciphering Soft Matter Morphology via Small-Angle Scattering with CREASE(-2D) with Arthi Jayaraman
Watch on YouTubeVideo summary
The webinar introduces CREASE(-2D), a computational reverse engineering analysis tool developed to interpret small-angle scattering data from soft matter systems where traditional analytical models fall short. Presented by Arthi Jayaraman, the method addresses complex challenges in analyzing polymers, colloids, and living nanoparticles that exhibit structural dispersity or unconventional shapes not covered by existing formulas. Unlike standard approaches that rely on fitting specific geometric forms like spheres or cylinders to one-dimensional averaged data, CREASE utilizes an evolutionary algorithm to explore a vast landscape of possible structural features simultaneously. This allows researchers to handle both isotropic systems with size distributions and anisotropic structures without losing critical information through azimuthal averaging, effectively bridging the gap between raw scattering patterns and real-space physical descriptions.
The core functionality of CREASE involves generating in silico replicas based on user-defined shapes and parameter ranges before training a machine learning model to rapidly predict scattering profiles. This pre-computation step accelerates the optimization process from days down to minutes on standard laptops, enabling near-real-time analysis during experiments. The tool outputs distributions of structural parameters—such as domain sizes, shapes, orientational order, and dispersity—rather than single values, acknowledging that multiple distinct structures can produce identical scattering patterns. To validate these results against experimental reality, the method incorporates additional physical properties like optical color or microscopy images; if the computed structure predicts a material property matching independent measurements, it confirms the structural interpretation is physically correct rather than merely mathematically consistent with the data.
The presentation highlights diverse applications ranging from bio-inspired melanin-mimicking superballs to floppy peptide tubes that change form under varying pH and temperature conditions. In cases involving anisotropic systems like twisted cellulose fibrils or sheared assemblies, CREASE successfully analyzes full two-dimensional scattering patterns to reveal hidden structural details such as tube eccentricity and tortuosity. The tool demonstrated its versatility by resolving ambiguous data where analytical models failed, identifying that what appeared to be simple tubes could actually represent a mixture of circular nanotubes and aggregated tape-like structures depending on environmental conditions. By combining computational speed with rigorous physical validation through property matching, CREASE provides researchers with robust distributions of structural features essential for understanding the multiscale organization in soft materials.
Read the full video transcript
maybe with a short introduction. First
of all, welcome everybody to the first
webinar of this series. Um,
so my name is Fedria Sebastian. I am
together with Yansko Ped which I see in
the audience and Santos from
uh we are putting together this webinar
series um in the context of the links
smile team which smile means of matter
in life. If you're not familiar with it
I'll share the link to the team so you
can look up and stay informed about the
upcoming seminar and activity within the
team. Um
so this is um the first uh seminar of
this series. I'm very happy to have yam
to present her work and very interesting
approach to analyzing data from smallus
scing
until now what is published about
polymers and my cells and so on but I
know that she's working on living nanop
particles so it's it's going to be I'm
sure very interesting to the audience
and to the team that we are interested
on so please arty I leave the word to
you.
>> Excellent. Okay, let me minimize this. I
won't be able to see anybody. Um, so
just if there's any questions or
anything like that, just speak up so I
can hear that there's someone online
because I'm going to minimize this
window if that's okay.
>> Yeah, that's fine.
>> I can keep an eye on the chat as well if
people see either of those. So, just
interrupt me as and when. No problem.
Okay. Um, thank you very much for the
invitation. Uh, good to see you, Yan. uh
uh for others who don't know many many
many years ago this age is both Yan and
myself when I was a posttock in 2006
um I read every single one of Yan's uh
review articles and papers on how to do
scattering calculations and and almost
20 years later here we are right so um
Yan you've had a very positive um
impression on my uh past academic career
so I'm happy to be here all All right.
Um, so I will talk about my former labs
work. Uh, I am not working on this topic
as of right now, but my former team
members are wrapping up uh, many of
their ongoing studies. I know some of
you in the audience are some of our
collaborators. So, nice to see you. I
saw the names. Um, and so what I will
share today was all work done uh, in the
last uh, whatever six to eight years at
University of Delaware with a lot of
recent work that I'll focus on. Um but
currently I have moved to IBM consulting
France uh where I get to play with the
computing tools for material science and
very much like the topic today it
involves solving problems uh finding
solutions to complex problems that other
methods are not able to do and that
theme still continues very much but now
the problems are more related to what
industry is doing it's still soft
materials I still continue to do
macroolelecules work so hope to interact
with you in that context as well all
right let me get my laser pointer
started. Uh there it is. Okay, great. So
um clearly this uh webinar is in the
context of folks who work with soft
materials who do scattering experiments
whether it's small angle X-ray
scattering, small angle neutron
scattering doesn't matter but that's the
audience I understand. uh so I want to
get directly to the idea of why small
angle scattering X-ray or neutron is
very valuable in the context of soft
materials and soft materials by that I
mean polymers colloids anything that's
not purely inorganic but rather organic
in the majority and there may be still
inorganic particles within uh uh the
blend with all these other materials so
that's what I mean by soft materials and
many of you who are here I will not
argue with me I'm sure that's this is a
very very powerful technique if you want
to understand spatial arrangements in
these materials at multiple scales.
While I'm not giving you motivation
today that I'm sure you know that most
soft materials function and application
comes from these multiscale structural
arrangements. So understanding how the
chains are organized, how the monomers
organize into the chain RG radius of
geration, how the chains organize and
then how the chains organized around
maybe nanoparticles. All of these
different level uh length scales of
structural arrangements is what
eventually gives you mechanical
properties, transport properties, uh
electronic uh maybe in some cases
thermal conductivity etc. So
understanding multi-length scale
structure is important and many of us
often go to small angle uh x-ray or
neutron scattering to understand
structural arrangements. Uh while I will
not talk about wide angle x-ray
scattering or wide angle neutron
scattering I've mostly seen wax uh that
is another technique that is also useful
when you have more precise order at the
atomistic scale and I won't be covering
that. It's there in the picture. So I
thought I should mention it that often
people also go towards wide angle X-ray
scattering in conjunction with small
angle to expand that length scale to go
even smaller with wide angle. But the
talk today focuses on problems which are
relying on small angle X-ray or neutron
scattering and the technique uh that
we've developed uh to interpret the
data. So what does the data look like
and what is the traditional analysis? Uh
many of you who have used small angle uh
scattering um have probably seen this
but if some of someone is new in the
audience uh usually the the output from
a small angle scattering experiment is
essentially the materials scattered
pattern uh for an incoming either X-ray
or neutron. The pattern itself is coming
about because of the chemical
composition of the material and the
structural arrangements of the material
as I said the multi-length scale
structural arrangement. So usually the
pattern is 2D where there's intensity of
scattered wave as a function of in this
case two wave vectors. But most people
don't always look at this 2D. I'm going
to call this 2D profile because in many
cases the structure is isotropic. And if
the structure is isotropic, you do not
have an anisotropic pattern that you
need to understand or interpret. So what
people do is they'll do a zutal angle
averaging. So imagine I average the
intensities along this arc that I'm
drawing. And so then I'll have high
intensity, low intensity, high
intensity, low intensity as I go
outwards, which then looks like this.
Intensity as a function of magnitude of
the wave vector. high values here as I
go farther. There are ripples and the
pattern of these ripples and the shape
is really a function of the arrangements
in your material. Okay. So when do we
care about these 2D patterns? We care
about these 2D patterns when there is
some information in my structure related
to anisotropy. This could be because the
structural arrangement is anisotropic or
there are anisotropic elements that have
local anisotropy but there is overall uh
isotropic arrangement. Regardless, some
of those signatures show up here and
then if we do this aimothal averaging,
we will lose that. So there is a need to
be able to handle 2D patterns with the
anisotropic patterns of uh present in
the scattering results because that
tells you about anisotropy. Second, if
suppose there is no anisotropy and your
structure is isotropic,
then when you have these 1D profiles in
most cases as we all have used many of
the Peterson models, many many models
exist out there that beautifully fit
this data and then once you fit that
model, you can get the parameters in the
model that you don't know. For this, of
course, you have to have expertise in a
knowing what the shapes probably are in
your material. two, uh, finding the
right model for that shape. And so if
those two are something you are able to
say yes and yes to, then you go fit, you
get the answers, problem solved. This is
what's been done for decades. But
sometimes you access structures either
through processing or some really cool
polymers formed by like my
collaborators. Then you lead to then you
cause structures in your system that may
not have an analytical model that exists
out there. And not all of us are Dr.
Peterson and so we can't go and develop
our own analytical model for every new
system. So what we've come up with is a
method here called crease that I'll
describe that a either handles this type
of problem where you're dealing with 1D
scattering profile but you don't have an
analogous beautiful analytical model
that is applicable for that system or
you have 2D and you don't want to do a
zuta averaging into 1D and you want to
analyze all of this. So this is the
expectation that we set out to have for
our method. I always say that no method
is perfect and the way we can describe a
method to be useful is if it satisfies
at the very least the expectations we
set out to satisfy. So our expectations
have to be reasonable. So here are the
three expectations that we set out to
do. They may be a bit ambitious but we
have accomplished these. And I also want
you to know what it doesn't do. So first
we can handle both X-ray and neutron
scattering small nuanced differences in
the way you handle the data and also in
the way you will do the algorithm as
I'll show you very small differences. So
we can handle both. For some systems we
can have both. The structures that we've
handled so far, our collaborators bring
us most of this data. So we're not doing
the experiments, but they come to us. In
many cases, these are unconventional
structures. Something does not fit
analytical models. So that's what we
want to be able to understand and
interpret. And then in cases where we
get 2D data, as I'll describe to you in
the latter half of my talk, we don't
want to do a zimaling. We want the
method to handle the full 2D profile
intensity as a function of wave vector
magnitude and azimutal angle full data
not as an image necessarily but it
should do the whole thing. Second I love
polymers. I think I've said this in
almost every one of the talks I've given
in my entire life since I learned about
polymers. Um and the beauty of it or the
challenges in it is the fact that there
is dispersity. What I mean by that is
everything is not precisely one
dimension. the chain length. So the
molecular weights tend to have
distributions. The arrangements in the
structure can have distributions even if
they're all the same shape. They may
have dimensional distributions or they
may even have different shapes of
arrangements. So we want to be able to
give that as an output. So not singular
value of certain parameters but rather
distribution of parameters that describe
our structure. The second part this is
also where I'll bring back the idea of
wax. If you remember I said that wide
angle X-ray scattering tends to work
really really well when you have
crystallin order or rather is very
useful when you have crystallin order
and that is not a problem we are trying
to understand we're not trying to
interpret the crystallin order either in
particles or in polymers but rather we
are focused on the amorphous or
disordered structures now the reason I
bring this up is you can imagine this as
trying to identify the real space
structure for a given scattering profile
right so So the math problem is slightly
different if you're trying to identify
one of the many ordered states the
system can access versus the
distribution of disordered states that
all fit my picture for the scattering
profile. So that's the math problem
difference. That doesn't mean we cannot
handle semi-crystallin polymers or
polymers and particles that have some
ordered arrangements. It's just it's the
type of interpretation that we do that
we should keep in mind. Okay, those are
the expectations. Let's see how well CRE
does. CRE stands for computational
reverse engineering analysis of
scattering experiments. The whole idea
is to be able to take as I said either
SAXS or SANS data from a system either
the 2D full 2D profile before a zenoil
averaging or 1D after a zeoil averaging
depending on your knowledge of the
system. So you should know when there's
anotropy when there isn't and you make
that decision. The method takes either
one or multiple profiles. So when I say
multiple, I want to be clear. It's the
same system. Maybe it has sachs and then
maybe it has some contrast matched s. So
then it's two pieces of information for
the same system or it might just be one
or the other. So that is taken as input.
Don't focus on the inside right now.
Let's look at the output. Crease's
intention is to provide the user
distributions of structural features. So
this is what this means is quantitative
description of structure in real space.
This might be domain sizes, domain
shapes, extent of orientational order,
dispersity in sizes, dispersed in shape.
Think of all of these physical
parameters that you tend to describe
your structures with and think of the
corresponding mathematical parameters
that describes it in real space. And
this is your choice because you know
your system. you know what you care
about in the structure. So it's these
mathematical descriptors whose values I
want to identify
and it would be great if we also have
visuals. So representative
three-dimensional structure structures
that describe some of these values of
structural features. So for some
represent representative cases we create
these 3D structures.
But keep in mind if you leave my talk at
the end of this thinking we only
generate 3D structures I have failed as
a speaker. So please keep in mind that
the main output is the distributions of
structural features. I also want you to
have an expectation very clear in your
mind. This is not a onetoone problem.
Meaning if I give a scattering profile I
have one answer one maybe distribution
of values. No, you may have multiple
answers where and you almost always have
multiple answers all of which give
similar scattering profile as the input
but then you have to put your physics
chemistry hat on to say which of these
answers is right. So multiple answers
for a given input and distributions of
answers and what we do inside is what
I'm going to describe in a bit more
detail. So basically our idea is we want
to reverse engineer all of the relevant
values of structural features whose
computed scattering matches the
experiment. Right? So it's an
optimization problem. Keep searching
through the values possible values of
those mathematical descriptors and
identify a family of values which all
give computed scattering that match up
the experiment. So that's the
optimization problem and we use an
evolutionary algorithm to do our
optimization. You may use something else
but the way it begins is we start the
first generation and now I'm using words
from the world of evolution
individuals and technically these
structural features are genes of those
individuals but if you want to use the
mathematical term I'll call these sets
sets of structural feature values and in
the previous case I gave you just an
example as domain shape and size. I
would like you to start thinking maybe
in the shape of a cylinder for example
because my next schematic is a cylinder.
So one could be the diameter of a
cylinder, length of a cylinder,
thickness, maybe some dispersity index
or something like that and a few more
structural features as you define it.
And I know there are analytical models
out there for cylinders. I'm just giving
this to you as a simple example. You'll
see the structures we interpret are way
more complex than cylinders. But this
helps me illustrate the method. So think
of this as diameter, length, thickness
and maybe dispersed steam one or more.
So I have values of those structural
features and these are sets of values
that I've started with and they're very
diverse sets in the beginning because I
know my values could be uh from A to B
for each of these structural features
where A to B is defined by the
scientist. It can't be a 100 nanometer
length cylinder. So maximum 100
nanometer it is definitely bigger than
10 nanometer 10 to 100 and so on. So I'd
have a diverse initial set where the
range of values I'm exploring is defined
by my scientist. And so for each of
these sets I need to calculate the
computed scattering. I'll come to how I
do that. I being my former team members
who did all of this work. How do we
calculate the computed scattering for
each of these is something coming up.
But we would do this for every set. How
many individuals you have in each
generation is your choice. How much
computing power do you have? How many
possible structures do you want to
simultaneously evaluate?
I will tell you we've come to a point
where these can be run on laptops. So
this should not be your limiting step.
You should be able to explore very
generously so you don't miss out on
right answers. Anyway, so in the first
generation once I have these individuals
and their computer scattering, I compare
each of these with the input comput in
input scattering and those that have a
good match are fit individuals or high
fitness individuals and those that do
not match well are essentially poor
fitness but they could still serve a
purpose of sharing their genes for the
next generation. So basically we start
from this initial generation and go
generation after generation after
generation by basically taking
individuals who are fitter and fitter
and fitter creating offsprings etc till
we go down to a final generation where
pretty much all of these individuals
computed scattering matches pretty well
with experiments and it's not really
changing much from the previous
generation to this. So fitness has
converged quite um uniformly among all
of the individuals. And so all these
sets of values are right answers. And we
obviously have to figure out which of
these makes physical sense. In many
cases they all make physical sense. In
some cases some may not. So basically
going from the top generation to the
bottom we're going from less fitness to
more fitness. And now we can look at
these final sets of answers to know all
possible answers that give a computed
scattering that matches experiment.
Okay, so now where how do we calculate
the computed scattering and where did
machine learning come to help us to
accelerate this? So originally when we
were doing this this goes back to all of
the knowledge I learned from the
Peterson papers um which is if you have
a shape a structure in this case as I
said let's take a simple example of a
cylinder I can place scatterers in the
cylinder and I can do a Dubai uh
scattering equation calculation which is
very well based in physics and so then I
can use that equation on each of these
pairs of scatterers and calculate the
icon. This is a pretty hefty calculation
and this is something that we can do
with computers but it takes a little bit
of time especially if now I have to do
this for each and every one of these
individuals and maybe the 100
generations or so that I might have to
explore. So this calculation starts to
really make it slow. So what we said was
we should be able to take the data of
the varying values of these structures
whatever they are. This could be an
amoeba shaped structure. It could be a
randomshaped structure. As long as you
have some information from another
analysis or characterization that says
it's amoeba shape. It could be any
shape. As long as you've mathematically
described it, we should now be able to
relate values of those mathematical
uh features to the corresponding ifq. We
would have this data. So let me go
through this a little bit better in more
detail so there's no confusion in your
mind. This is how we would get started.
We would have to decide on the shape of
the structure, right? That's what our
domain or material science expertise
told us. The shapes look like this.
Maybe you're thinking amoeba. Feel
creative. I'm going to use a simple
cylinder as an example. We make using a
computing a code static structures. So
keep in mind there is no molecular
simulation here even though that's one
of my favorite tools. Here we're just
creating static structures with codes
that allow us to vary values of these
structural features. And for each
structure I place scatterers which is
point scatterers delta functions and I
can calculate scattering profiles. If I
do this repeatedly for various values of
dt thickness and length I basically have
a database of these values and its
corresponding scattering profile. So now
with that database I can train a machine
learning model to relate values of the
DTL directly to icon. This means I don't
have to do that hefty Dubai equation in
the genetic algorithm loop and instead I
can use this machine learning model
right there for every set I take the
values of that set put it in the machine
learning model generate its icon very
fast training was done outside of this
loop. So if I do this entire genetic
algorithm loop, now we went from the
slow early version of 1 to two days to
less than an hour to a few minutes
depending on your laptop. So it's
basically allowed us to now create a
tool that can be analyzing while you're
also making measurements and it's done
fast as long as the training has been
done before. If you have a clear picture
of your shape and um the ranges of
values of the structural features for
those shapes, the training can be done
in one to two days at most. And then
this run is very fast on a laptop. Okay.
So what have we done so far with crease?
I don't have plenty of time. So I've
decided to share only two problems with
you. Um but I wanted to share the
breadth of problems to which we have
applied crease and and you can see the
variety of shapes and I'll try to
highlight specific challenges we had in
each of these problems which led us to
use crease. So originally we started
this problem for a project with Karen
Woolly and Darren Pochan. Let me just
make sure I'm okay on time. Yeah. And
Karen Woolly's lab is incredibly
talented in making some exotic polymers.
Let me not follow polymer physics
Gaussian statistics for their
confirmations. So even though the
structures they were forming which we
could see in microscopy looked my poor
shell spherical myels for which we have
analytical models because these chains
were very um exotic. they were not
following Gaussian statistics and the
analytical models fits didn't make any
sense. So we came up with crease and
then we basically analyzed what types of
shapes and structures for the shell and
the core we could explore and the
dispersity etc. This is where it all
started and because we don't deal with
molecules and we don't leave deal with
chains we were not limited in handling
only certain number of chain
confirmations or certain arrangements.
we could explore many more dimensions as
long as it was um it was something that
we explored within genetic algorithm.
Similarly, we also used the very exotic
polymers that they made. These are poly
glucose carbonate backbone and charged
side chains. So, persistence length of
two types. The most the the the very
exciting polymers. Don't ask me to say
the chemistry of these polymers. You'll
have to look at the paper. Anyway, so
other shapes and sometimes we don't know
the shape very well from the microscopy.
So then we test multiple shapes with
their corresponding structural features.
Then we move to something that is close
to your team, the lipids team, the smile
team at links, which is essentially
vesicles. Here it's not lipid vesicles.
These were protein peptide based
vesicles from my colleague late Christy
kick. uh her lab was interested in
understanding vesicle dimensions and
they had a lot of poly dispersity. So
here the poly dispersity is what made
this problem hard to fit with analytical
models. If you removed poly dispersity
and you had mono disperse vesicles there
are great analytical models out there
but here she was having lots of
dispersity in the dimensions of the
vesicles. We also looked at a system
from Tim Lodge and Frank Bates which are
metal cellulose fibrals. a metal
cellulose or cellulose derivatives that
form twisted fibrals and a lot of the
science that they got by fitting
analytical models were fascinating and
the results were just surprising. So one
of the things we wanted to do was
challenge and say is it because of the
choice of the analytical model or is
this the same result we would get if we
did crease. The answer in this case was
we got the same result. Um and it's okay
because here we were just testing if
this was a limitation. The answer was a
limitation of the analytical model. It
was not. Those were systems where
molecules formed forms
and that form factor was what we cared
about. But the structural arrangement
because dilute it didn't matter. Now we
looked at systems where the form the
example I'll show you is where the form
is pretty straightforward. Spherical
particles which has dispersity but the
structural arrangement is what we want
to go after. So I'll highlight that. I
won't go more into it right now. We also
looked at work with Bhnesh Parati where
we focused on both changing form
changing structure as we changed
conditions. So we didn't assume did not
assume a form and then interpret
structure. Instead we were interpreting
both simultaneously which is the holy
grail with an at least isotropic um soft
material systems. The last part will be
an example. So there'll be two examples
today. One from the previous slide and
one from this slide where we've expanded
to uh apply crease to systems that have
anisotropy. The anisotropy could be in
the particles, could be in the
arrangements, could be in the self
assembled system upon shearing. It
doesn't matter what gave rise to the
anisotropy, but crease can handle 2D
scattering profiles that are relevant
when you're doing anotropic structures.
So, I will highlight one example from
here and I saw Simona in the audience.
So Simona will find this very very
familiar. Okay. So first let's get
started with a system where my form so
my particles are soft particles uh and
they have a form which is spherical and
I'll show you proof that it is spherical
and uh what we want to go after is the
arrangement. So it's a mixture of
nanoparticles and I want to understand
degree of mixing and also mixture
composition. just because of the way the
part the the assembled state is made.
You cannot guess the mixture of the
composition. You may know the mixture of
the initial uh solution but not the self
assembled particle. So this is what we
want to find out. These are both related
to structural arrangements. So a little
bit more about this project. This is uh
this was supported by air force when I
was in the US and our collaborators who
worked very very closely in so many
publications with us and created this
gathering data. Ali noa's lab, Nathan's
lab did the synthesis and so we were a
collaborative team. Okay. So what is the
system? It is a bio inpired system and I
know that's something of interest to the
smile lab soft matter in life. Um and
here the motivation was to mimic
melanin. All of us have seen how we
react to very sunny days how our skin
changes color. So this project was all
about mimicking what organisms of every
form do to create color in their skin
and create synthetic materials that did
the same. So one of the ways we were
accomplishing that and this the whole
phenomena is called structural color. So
achieving that is through controlling
structure and the chemistry of the
particles. So we had two types of
particles silica and polydopamine which
is the melanin mimic and we created
mixtures of those we being Ali
Dinojala's lab and an unveil um where
they basically started with a certain
composition of the particles h in one
solvent and then they blended in another
solvent. So in this case octanol was
already there water was added that led
to this reverse emulsion formation and
then basically the droplets lost the
water outside and as the water left the
droplet you form these self assembled
superbles and so what we wanted to know
was what is the arrangement within the
super bowl so the degree of mixing and
the mixture composition in the super
bowls so for this uh Ali's lab did sachs
and sands when you have mixtures it's
nice to contrast match one against the
other the solvent. So you get one of the
particles highlighted and then SAX gave
the idea of the whole particle
especially when you have a pure silica
or pure polyopamine particle. So both
profiles were used. I also want to
motivate as I said why we care about the
structural aspects like mixing and
silica composition because that's what
gives me color and at the end of the day
this was what the project was about is
to create color in synthetic materials
by mimicking melon. So let's now go
directly to scattering results. Just
very briefly before I say that before we
apply crease to any experimental data my
team members had always created insilico
replicas for which they knew the answers
and when they put the scattering profile
from that incilico system and ran it
through crease they could compare the
output to the answers they knew were
correct. This is the way we knew crease
worked well on pretty much perfect
systems before we applied it to the
experimental data. So that we know the
method works well and now let's see how
it works on experimental data. Like I
said if you have sands and you have
smearing you have to bring in you can do
that in the computed scattering
calculation. Okay. So these particles
sacks on them shows they are spherical
particles both in microscopy and the
form factor fit is perfect. you get the
diameter and distribution. So it is poly
disperse. And then when you look at the
super balls in the microscopy, you see
that there's arrangement of the
particles that is not ordered. So
amorphous arrangements. And now this is
the sax data from pure melanin and pure
silica. So think of the composition as
100% silica. Here 0% here. As I said, we
had a whole series of intermediate
compositions that were handled with
sands. Um and what you see here is if I
try to fit this with a sticky hard
sphere model which is uh quite often
used in such scenarios I do not get a
good fit. Same here did not get a good
fit. But if we did the final generations
computed scattering on top of the
experiment you see there's no difference
which means they're perfect match. Just
because it's a match doesn't mean the
answers are right. So the only way we
will know that the final answers that
come out of crease are correct is if
they also computationally predict the
same color as we see in experiments
because in this case color is coming
from the structural features. So if you
have other properties that are coming
because of structure, you could use that
property measurement, compare that to
the calculated property from the 3D
representations and that can be a way to
know if your answers are right or if
something is numerically right in the
genetic algorithm but physically not
right. So here we collected color sorry
computed color on the representative 3D
structures we got out of crease. So the
synthesis of the materials happened.
They went for experimental color
response measurement. These materials
went for scattering X-ray and neutron
depending on mixtures and pure. Uh the
scattering profiles like I showed you
came to crease. priest did the
calculations gave the structural
features distributions as well as
representative 3D structures that when
used in you know optical color
calculations
essentially gives a computed color
response you can compare the two
together if the answers are physically
correct then these two should match and
it's not one color response versus one
experimental prof response as you can
see multiple samples from the same
materials were tested here. Multiple
right answers color calculation together
gave the error bar. So error bar sources
are different but within error they have
exactly the same shape and exactly the
same spot in the color map. This means
that the final answers are correct
because they produce the same color as
an experiments that comes from
structure.
Like I said that was just for structure.
Now imagine if my particles were
changing form as I changed concentration
or temperature. That's what we did in
the paper with Bubesh Bari where
basically um silica coated with
surfactants was changing form as you
changed pH and temperature. So you
couldn't just assume oh it's a sphere
sphere. Let's just focus on structure.
We had to do both together. Okay. So
moving from isotropic systems to now
anisotropic. uh I want to show you our
capability that we have so far with
crease 2D and I want to highlight again
that before the experimental systems
I'll talk about today we had to do all
of this in silico where we had
structures whose answers we knew we
created these structures the machine
learning model did not see it so when we
input that known structure scattering
profile to the crease algorithm it
produces mathematical values and
distributions and that can then be
compared with the right answer. So this
is how we first demonstrated we can
handle entire 2D scattering patterns and
get right answers. So once we had that
capability then we were lucky enough to
collaborate with Dave Adams at
University of Glasgow.
His group has been focused for many
years on very creative assembly of
deptide systems in variety of solvents
and salts because peptides have charges
on them. They can respond to salt. They
have interactions that are very unique
with water and other solvents,
hydrophobicity, hydrophilicity. So with
a combination of deptides and their
isomers and salt and solvent, they were
able to form variety of assembled
structures. For some systems, these
assembled structures are very easy to
understand from microscopy, cryotm, etc.
But for certain uh deptides, it was not
so straightforward as you will see that
this their their um their scattering
data didn't fit analytical models for
tubes or cylinders. And so they came to
us with we have some anisotropy we think
and even in the isotropic systems we're
not able to understand why the
analytical models don't give a good fit.
I also want to highlight how the data
looks. Um I usually emphasize that this
data is ugly but I know Dave student is
here in former students here in the
audience. So when I say ugly data I want
to be clear. I'm not saying that the
data is bad. The data is beautiful.
But that is not perfect. And so when we
do the computed scattering calculations,
we have to impose such imperfections on
the computed scattering so that you can
fairly compare it with experiments.
So here are just four of many many many
um profiles that they have um gathered
from various synretron sources. And so
I'll highlight two systems uh today. But
before I highlight that, I have to tell
you that our interactions with Dave's
lab as well as ourselves, we were first
trying to find what were the right night
right types of structural features that
describes the system. And after much
back and forth, we came to a point where
we realized what was really happening
was these dieptide tubes that were
forming were very floppy. This is our um
interpretation. And these floppy tubes
have cross-sections that have
variability with along the tube. So
think of it like a straw that you would
never drink out of because it's so
floppy, right? And so the parameters
that describe these are obviously
dimensions. But in particular, I want
you to look at this parameter called
eccentricity that says a cross-section
is circular eccentricity one. If it's
tapelike, eccentricity, sorry,
eccentricity zero. Tapeelike is
eccentricity one. and anything in
between tells you that it's not circular
and not tape like and some.
Um the left column here describes all
that we wanted to have represented in
our structural features but these are
written out in English in physics
chemistry language and these are all the
mathematical parameters that allow us to
interpret these aspects of orientational
order tortuosity stuffies. So when I
describe the answers to you, they'll be
in this format. But I'll do the
interpretation to the physical
interpretation with my explanations. I
also want to highlight uh two things.
One is the choice of the machine
learning model is really yours. It's how
much data you collect from those static
structural variations that you create.
3,00 2,000 is a reasonable data set. XG
boost model does well. I also want to
highlight that we do not deal with 2D
scattering patterns as images. Instead,
we deal with this as a 2D vector of
intensity as a function of ray vector
magnitude and a zig line. And you can
work with polar coordinates or cartisian
coordinates. It's totally up to you.
It's back and forth. Conversion is
straightforward. Okay. So, this is the
data. And now we go towards looking at
the analysis. Okay. So when we give this
input we see there's no anisotropy.
So my values of kappa which describe the
orientational order and in this case
it's in log scale should be very very
small and it is the plot here is a
violin plot of all possible structural
features and the values that crease says
the final generation has. It doesn't
tell you sets. There is a set could be a
value d here, omega here, alpha here,
maybe epsilon over there, kappa here. So
it just has a collection of all the
sets, but it doesn't look at the sets
individually. But we can look at the
sets individually. And so here you can
see three families of answers. You can
do clustering of the final answers into
families. And you can see the different
parameters and their values. So like I
said, I'll do the conversion to physical
language. Diameter 100 anstroms, which I
believe Dave's lab has seen many times
with their tubes or whatever they see in
the microscopy. These are tapes because
the eccentricity is one. There is some
variability in the eccentricity. We do
not pay attention to these values
because they do not the eccentricity
dispersity does not have a unique effect
on the scattering profile. So it can
give you many but the presence of
dispersity is there. This we can trust
the mean value of eccentricity as we've
shown in the past with our incilical
structures. Kappa and the log scale is
very small. Now let's look at the three
representative structures. They have
some tortuosity. There's no anisotropic
order. Orientational order is missing.
These are just floppy tubes. Sorry,
floppy tapes.
Now let's look at this other input
scattering profile. again family of
answers. You can see now the kappa has
moved up and it's in log scale which
means that there is anisotropy and you
see it and so now these are the
reconstructed computed scattering before
we impose this this
flaws that we find in the experimental
profile. Now we have two sets of
families of answers which differ in
their story. One says it's 100 angstroms
circular cross-section. So these are
tubes
high anisotropic order. So orientational
order is high tortuosity is not very
high. One is now 300 angstroms tapes
and yes there's some anisotropy slightly
less some torch. Okay. So how can this
be two sets of answers? Is one wrong?
When we proposed this to Dave, he
suggested that actually there is merit
to both because there can be tubes that
are 100 anstroms. But salt interaction,
salt induced aggregation or solvent
induced aggregation can lead to tubes
aggregating next to each other looking
like tapes that can have larger
diameters. So we believe that there is
merit to this picture of 100 angstrom
tubes or 300 angstrom tape like
structures existing simultaneously.
So how do we ever know that the crease
interpretation is right? One is exactly
what I showed you before by property
calculation. One would be molecular
dynamic simulations if you have force
fields but you'll never get to all the
length scales. It's going to be very
very very hefty calculations. You only
do that for the select few for which you
have a good force field. The other is
microscopy images and this is the
knowledge that Dave had that enabled us
to say that this what we observed with
these two families of answers is not
wrong. All right. So I'm going to stop
there. I have one more slide and I'll
come back to the acknowledgement slide.
I know this team here smiles team is
very interested in the lipid uh vesicle
crease story and yes we do have a story
that my former postto Rohan has
completed with folks in illoba Leonel
Porsar and Ursula Perez this is
manuscript that'll be submitted very
soon uh and then there's one more
manuscript that my former postto Andrea
Finley is wrapping up this is with Dave
and Simona on using crease 2D to analyze
realology and SAS happening
simultaneously. Okay, so thanks to the
members who did all the work. Uh thanks
to the experimentalist who gave us a ton
of different forms of data and trusted
us to to to use crease in a in in a way
to give them interpretation. All the
codes are in the crease GitHub page on
my lab. So if you look at RTJ Ram on
GitHub, you'll find it past financial
support and thank you for your
attention.