CPHR Seminar Series - Alison Motsinger-Reif
Watch on YouTubeVideo summary
Alison Motsinger-Reif, Chief of the Bioinformatics and Computational Biology Branch at NIEHS, presented an in-depth overview of the Personalized Environment and Genes (PEGS) study and its integration with the All of Us ancillary research. The PEGS cohort comprises approximately 20,000 individuals in North Carolina who have undergone genome sequencing and provided extensive data through three massive surveys containing over 1,700 questions covering clinical conditions, lifestyle factors, diet, sleep, and residential or occupational histories. By utilizing geospatial linkages, the study maps participant addresses against environmental hazards, air pollution, land use, and climate data across 99 of North Carolina's counties, creating a robust foundation for analyzing the complex interplay between genetics and the environment.
The presentation highlighted several key research themes, including Exposome-Wide Association Studies (XWAS) that identified significant links between blood type A Rh-negative and heart attacks, as well as paternal education levels and cardiovascular outcomes. Researchers also found strong effects of air pollution mixtures on autoimmune skin diseases comparable to smoking history, while population-specific analyses revealed different HLA gene classes associated with asthma in African American versus European ancestry participants. Furthermore, despite small sample sizes, genome-wide significant hits were identified for rare outcomes like gestational hypertension, and poly exposure scores aggregating environmental factors demonstrated superior predictive power for Type 2 diabetes compared to traditional polygenic risk scores, even when combined with clinical and genetic data.
To support these findings, the team developed "PEGS Explorer," an interactive web tool that allows users to explore associations, stratify by covariates, and download data, while the study is transitioning from cross-sectional prevalence data to longitudinal insights through linkages with electronic health records. The scope of measurable metabolomic data includes host metabolites, dietary components, environmental pollutants, and pharmaceuticals, which are integrated with genomic, epigenetic, and social determinants of health to conduct risk score analyses. Addressing recruitment challenges in rural areas, the study successfully partnered with trusted community institutions like churches and utilized women's health initiatives, electronic follow-ups, and incentives, while serving as a core facility for clinical procedures such as skin biopsies and inhalation assays.
Looking toward the future, the All of Us ancillary study aims to integrate deep exposomics via mass spectrometry and biospecimens to investigate incident diabetes and complications, leveraging upcoming collaborations with NIH Common Fund projects. Motsinger-Reif emphasized that while moving from association to causality in exposomics is more difficult than in genetics due to the complexity of harmonizing environmental data, current efforts focus on hypothesis generation and screening strategies rather than immediate clinical application. She called for greater openness in data sharing and interdisciplinary collaboration involving basic biologists, mechanistic biologists, and toxicologists to further disentangle correlations between poly-exposure scores and poly-social scores, ultimately advancing the understanding of how environmental factors interact with human biology.
Read the full video transcript
welcome
everyone I appreciate everyone that's
joined us today both uh in person and
online uh today we have a outside uh
guest to speak but from the NIH I'm from
our North Carolina Branch at National
Institute for Environmental Health
Sciences aliser Allison moninger rif
will be joining us and presenting her
work um from nihs and and uh she's done
a lot of work at the intersection of
genomics environment and a number of
different variables she is c-i of a
cohort called pegs which I should have
actually looked what this stood for but
she probably going to mention it and
it's an environmental uh collection of
about 20,000 people in the North
Carolina area that also has genome
sequence data it has a deep
environmental exposure data and she's
looked at that data in a number of
different ways that she'll talk about
she's also leading a um a study that is
an anary study to all of us which he'll
talk about which is gathering uh deep uh
deep um expose them data to link with
all the other types of data we have to
look at the risk of incident type two
diabetes in the environment she is a
let's see your uh their chief of your
branch which is the uh bio statistics
and computational biology Branch at n
niehs and been a great collaborator in
these projects so Allison I'll turn it
over to you thank you so much thank you
Josh for the invitation um thank you
everybody for their time especially as
many things are going on right now I
nobody's everybody's time and attention
is even scarcer than normal so I'm
really grateful um to get this time with
you guys um I'm going to attempt to
share my screen now and hopefully we'll
make that transition just
fine right do you guys see one big
slide excellent love it when the magic
works um so as Josh said I'm a human
geneticist um at
nihs um I come from a human genetics
background um Josh and I are both vendy
alums um at different times so I think
I'm supposed to say anchor down to him
um but that that slang uh is after when
I actually left um and so it's it's been
really fun like I said coming from a
human genetics perspective I was a
faculty member at NC State for a long
time in their statistics department and
B informatics Research Center and then I
joined in IHS about 6 years ago
um right now for the leadership role in
the branch and have had the wonderful
opportunity um to work with this study
um that Josh mentioned pegs stands for
the personalized environment and genes
study and I'll tell you a little bit
more um about it as we go
along so overall um I've gotten away
with in my career sort of a mix of
methods developments um from a
statistical biostatistical computational
biology um point of view as well as
really doing a lot of Applied work in
the the the balance of my portfolio has
changed Through The Years um but really
enjoying doing both of those things um
I'm not that clever a statistician I'm
not going to be in my office deriving
the asymptotics to anything new I steal
all my methods um challenges from the
real data um that I'm working on so I'm
always really transparent and I guess if
you highlight some themes I've worked in
machine learning pharmacogenomics
thinking about Gene Gene and Gene
environment interactions dose response
modeling and really I'll deal with
whatever m is huge and popular and um
and challenging so so this data I'm
going to talk about now has been a real
fun way to to really think about diverse
data um I hope it's okay um I'm going to
take you a little bit on a whirlwind
tour um here and appreciate your
patience I'm going to tell you some of
the things we've been doing over the
last two to three years um with this
study that I've mentioned um from a
number of different angles um sort of
the the studies I'm going to talk about
are prioritized based on mostly fellows
and trainees interest and where they
want to go in their career as well as
what we think um is going to be
scientifically um impactful and as you
mentioned I'm going to talk first about
the work we've been doing in pegs and
then we'll end the discussion with some
of the work that's ongoing in the the
partner research study um that Josh
mentioned with all of us and I hope
you'll be able to see the connection
between some of the kinds of work and
the way we've been thinking about things
um in the pegs data and how that um Can
can be really applicable in the types of
data that that all of us is now really
thinking about about collecting um that
I think is is really exciting so I hope
you're you're willing to hang with me um
we'll we'll move around um across topics
um so I'm going to use the the phrase we
um very generously right like any good
Pi I actually don't do anything um
everything that I present as we um is
from a really talented um group of folks
um that have done the the bulk of this
work Dr Freda OCT is a um contractor
staff scientist Dylan Lloyd is a grad
student that just graduated and we'll be
moving to Harvard next month for a
postto John house is a a staff scientist
Jasmine Mack's another graduate student
that'll defend this summer and then Ryan
Campbell is another um contract staff
member and Joe is a is a postto so
really I'm going to take credit for all
their work as well as I have to go on
and highlight other key
co-conspirators um that help me get away
um with all of this fun stuff um the
folks collaborating in pegs Jan Hall has
been the co-pi of pegs so she's retiring
next month Dave Fargo Charles Smith and
Messier both are are all fantastic
bioinformaticist computational
scientists um tenure track investigators
then I'll highlight Jeff and Josh and
others at all of us um that have really
been the the co-conspirators on the
ancillary study Dr Rick week our
Institute director and one-on-one
collaborator um at niehs that has
supported all this work um as
well so I'm just going to sort of
introduce here I don't I don't know what
my audience is on the mix of geneticists
versus other interdisciplinary
scientists but I want to sort of cue up
what we're doing within sort of an
exposomics framework um so like I said I
was trained initially as a geneticist so
right I'm certainly um used to thinking
about what are the what's the genetic
ideology of common complex traits um but
as I did more pharmacogenomics and
moving more and more into Environmental
Health Sciences um really am
appreciating how much more information
there is to gather from an exposomics um
standpoint and so when I say exposomics
I'm going to be really really inclusive
in that so here I mean things like that
you probably think about as
environmental exposures what's your
physical environment how much air
pollution is near your house or near
your work what sort of chemicals um do
you use in your job occupationally every
day so so really broadly but I also want
to make sure we include in that the
social environment whether it's social
determinance of Health um or other um
sorts of of social factors um that that
control the the body's physiological
response um things that you might
consider lifestyle diet exercise sleep
that kind of thing and then also within
exposomics the the body's response so
how do does the body manage that whether
that's exposomics metabolomics right
what how does the body process um these
chemicals that they're exposed to um Etc
and hopefully you'll see I I'm very
inclusive in that and sort of an
exposomics approach exposomics data to
me and the way I'm going to use it can
come from multiple sources so that could
be things like survey questions that I'm
going to talk a lot about or that could
come from geospatial estimates of
exposure right how do we we've got
estimates of air pollution or other
sorts of
exposures that could um include you know
sort of omix data right metabolomics
exposomics um as a as a as a proxy for
your environmental exposure so try to
think about that that really
inclusively so now hopefully you're sort
of queuing up to the framework and I'll
give you some more details on the peg
study um and and talked about what we're
doing with it so as Josh mentioned this
is It's a relatively small cohort
especially from a genomics um
perspective but it's a decently large
cohort from an environmental health
sciences perspective and we've spent the
last five or six years really Gathering
these multiple data types um in the
study so we've got a lot of of data on
genetics um over the years this cohort
has been around for almost 20 years and
there's been some halfhazard sort of
candidate Gene collection along the way
um but a major effort a few years ago to
collect hold genome sequencing data on
everybody we could afford at the time
which is about 4700 individuals we
recently last year we able to get um
methylation data epigenetic data from
the Epic chip on everybody that we have
whole genome sequencing um for
phenotypes um we have a number of
self-reported diseases and conditions um
from our questionnaires and I'm going to
tag in more detail on our questionnaires
when I talk about the environmental
exposure so pegs has huge questionnaires
patently absurdly huge questionnaires
compared to to what's often collected so
we have three questionnaires one we call
Health and exposure um and this gets to
a lot of has a doctor ever diagnosed you
with a b or c and then in your work or
in your home have you ever been exposed
to X Y or z a lot of questions stolen
from well-established instruments
studies like inhes nurses Health Etc and
then we have two other surveys that were
taken by fewer individuals but really
extensive on what we call our internal
and external exposome surveys so the
internal exposome you could probably
relabel lifestyle and that fit pretty
well so what medications are you taking
what physical activity do you
participate in um that's sleep um stress
that sort of thing then the external
exposome has a lot of detailed
residential and occupational um
exposures and then we also have a
continuously growing um and really
dynamically growing right now um set of
Expos meur es from geospatial linkages
so we have addresses and address
histories for all of our participants in
pegs including things like the longest
lived childhood and the longest lived
adult address where we've linked a
growing number of exposure
estimates just to give you a snapshot of
some of the demographics of pegs the
average age when they took the health
and exposure um survey is about 50 we've
really got some pretty broad diversity
in terms of Education levels you can see
some graphs um there um we are mostly a
Caucasian um ancestry cohort about 60%
but almost 28% come from africanamerican
ancestry which is which is decently
unique actually um we've got more women
than we do men it's a volunteer cohort
that happens um pretty steadily and then
we do actually have a pretty good
diverse range of sort of incomes um and
other um sort of measures of
socioeconomic
status just to give you a couple
snapshots on those different surveys um
so in that health and exposure um survey
we asked questions about over 120 sort
of clinical conditions and disease
States we've got some snapshots here um
you can imagine we have sort of about
population prevalence um of most of the
common complex diseases were slightly
healthier than the North Carolina
population um but but not by much um and
like said to give you um just a snapshot
of some of those those lifestyle um
factors as well then I mentioned these
exposome surveys are external and
internal exposome um over 200 questions
each exposome has things like the
characteristics of your current and past
homes workplace characteristics chemical
and metalic exposures hobby exposures
nuv light exposures and then internal
exposome has medications vam vitamins
supplements uh drug treatments um
physical activities stress infection
sleep diet um s sibling twin birth order
um genetic history in total between
those three surveys our our participants
actually answer about 1,700 questions um
which is anybody's everever been
involved in sort of that that design
that's that's a lot so it's really
really rich survey questions then I
mentioned our GIS data we've got a
growing number of data layers there our
first linkages were from to what I've
learned is actually a technical term
environmental bads so it's how far away
do you live from these environmental
bads things like airports Kos are caged
animal feeding operations that are not
entirely unique but a North Carolina
specific um exposure and then things
like cell towers drinking water dry
cleaners hazardous wa sites um you can
see the list um here I've just said
distance to these um environmental bads
um for again all of our addresses now
and longest lived um adult and then
longest lived childhood if that data is
available the gis data we've got a
growing number of linkages so we've done
a lot of linkages to measures of
neighborhood socio economics and
structure um a lot to air pollution
whether that's wildfires or
pm2.5 again a growing number of those
point source metrics distance to
environmental bads um we're working on
more and more particulate matter
linkages um a lot of high resolution
land use and land covers um measurements
things like green space Urban density um
a number of climate and climate change
um measures including temperature
extremes and and seasonal um variability
and and a lot of sort of linking with
with public open data API sorts of
things we have great coverage over the
the entire state of North Carolina we
definitely are biased to to sort of the
Research Triangle Park area here where
NS is located but we have participants
in 99 out of the 100 counties in the
state of North Carolina I also am from
North Carolina this is where I grew up
um so it has lots of sort of personal
resonance um for me now I'm going to
start this Whirlwind tour of some of the
things um that we have worked on and
thought about um so I I hope this you
know helps show the the kinds of things
we're we're thinking about so former
proo Dr Unice Lee um LED some work she's
really passionate about cardiopulmonary
um outcomes and wants to do that um for
her career and led what we call an exos
on sort of atherogenic cardiovascular
disease outcomes so exas meaning
exposome wide Association study you can
think of this as analogous to a Goos if
you're familiar with that area but we're
really in a sort of a hypothesis
Generation Um and sort of Big Data
exploration space so she took all of our
exposures from the questionnaires that
over 1700 and did Association analysis
where she tested you know exposure by
exposure with association with one of
several atherogenic cardiovascular um
outcomes I'm showing you here a tile
plot um so on the the bottom x axis you
see the the exposures that that question
on the Y AIS you see the cardiovascular
outcome that's associated with and
everything that has uh a colored thing
on this tiop plot had a significant
Association after a pretty rigid FDR
multiple testing control the top
actually has the effect sizes um of
those and the bottom just has it
magnified by by the strength of
Association um so doing this um found
some some interesting results we were
able to reproduce a lot of things that
make us feel sort of solid about the
cohort and the results we're finding um
and found some new things that hadn't
been reported um before in particular
strong associations with blood type um a
Rh negative so a negative with heart
attack paint related exposures with
stroke biohazardous materials um at work
with arhythmia and strong associations
with paternal education level and
several of those those outcomes so you
can start to see we can start explore
what are shared environmental
associations and what are unique um and
sort of look across that space you'll
see that we've scaled up this xwat
approach to multiple outcomes um
following her work we've also done some
very targeted sort of Gene environment
analysis work um so here we we worked
with canidate genes so we were
interested in what's the association of
the sort of distance to these Kos we
call them these caged animal feeding
operations which are really really
they're they're what they sound like
they're they're very dense industrial
agricultural um sites that have
absolutely terrible um environmental
exposures there's a a town to the east
um here called Smithfield that when you
drive through it you smell the exposure
for the Kayo um that is there your tires
will smell like that Kayo even when you
get home to the garage and Raleigh you
have to spray them off if you drove
through
Smithfield um so these really are are
pretty dramatic sort of exposures so
we're interested in associations with IM
immune mediated diseases and that
distance um to these caged animal
feeding operations in particular looking
at um Gene environment interactions with
known genes that are associated to these
these disorders so we looked at
individual diseases as well as grouped
diseases um and what I'm showing you
here is a pretty dramatic example of a
gene environment interaction so on the
y- AIS here you've got the distance to
the Kos so this is how far away our
participants live from one of those um
kfo facilities our y has the odds ratio
so you're seeing our effect size um and
here I'm showing you the the sort of
what the environmental effect is um
alone and then I'm showing you that
environmental effect in odds ratio by
genotype for individuals with this um
the homozygous minor alal you see a very
very strong um interaction effect um so
we're interested in report these on
multiple sort of genes in
ptpn22 um ahr pathway Gene interactions
with with Kos um exposures that were
really um
convincing so like at G by E we're also
interested in those geospatial um
linkages that I mentioned to you and uh
Kyle Messier that I mentioned there is
really an expert in developing the the
actual estimates um of exposure um so
right here we're looking at air
pollution you can imagine there this
this data comes from different EPA
monitoring sites and there's not an even
distribution of of those monitors so we
have to fit and um estimate these
different components of air pollution
and you can see a bunch of North
Carolina maps um with the values um for
those different components of um air
pollution and he takes really
sophisticated beian modeling sort of
approaches to associate those air
pollution um measures with um sort of in
this case the the traine was really
interested in autoimmune skin diseases
and long story short found some very
strong associations when you looked at
the mixtures so if you looked at these
individual components you didn't see the
effect but when you're able to jointly
model um the different components of
small um particulate matter um air
pollution was was able to see this and
to to give you a little bit of an anchor
um I said I know I'm going through it
quickly but sort of the strength of
these associations were about the same
magnitude of whether or not you had a
history as a smoker which comes up in
autoimmune diseases and many other um
common complex traits um all the time
this really is a strong magnitude um of
effect like said on par with whether or
not you
smoke additionally we mentioned we've
got whole genome sequencing um and so
we've been able to call the hlaa MHC
region um to six-digit resolution and
again this is work from Unice Lee the
the woman that was doing the xos studies
and did MHC region mapping with adult
onset um asthma so she tested population
specific specific snip and HLA alal and
then protein um associations and did it
um stratified for the two different
genetic ancestry groups um that we've
got here um and long story short she
found some really interesting sort of
population specific effects um so really
found that in um sort of our
participants with um African-American
ancestry is really H and HLA do o I'm
I'm sorry reverse that um in our
European subjects it was H and H O
and that class one HLA La genes were
strongly associated with our
participants of the European ancestry
and it was actually class two genes for
participants with African ancestry um
which is interesting and and we're we're
building on a lot of sort of the the
pipelines and work um she developed for
us for that HLA MHC analyses for for
stuff that we interested in in all of us
cohort um as
well um another really talented graduate
student um wanted to to pursue um gwas
analysis with adverse pregnancy outcomes
like I said I'm a human geneticist I was
incredibly skeptical about sort of the
the power um of this study to do any gws
mapping and especially in sort of rare
outcomes like adverse pregnancy outcomes
but it's one of those moments I'm glad I
didn't listen to me and I listened um to
her which happens to me all the time um
but she was really interested in that um
so did GS mapping in pegs of gestational
hypertension you can see we've got a
quite a small um sample size um when you
consider sort of what you're used to for
gwos mapping um but here you can see
that Manhattan plot down to the left
we've got a genome wide significant hit
and a couple of genes and including this
um rare B Gene it's a retinoic acid
signaling Gene um we were able to
replicate that in the UK um biobank you
can see there um sort of the the lead
snip there um and really interesting
rare be makes a ton of biological sense
so it's um been studied and shown to be
Associated um with sort of severity of
protura um and and in preeclamptic
studies um particularly it come up in
the fetal genome but we're seeing it
here now in the maternal genome um so
that paper was recently published um and
she's going to continue um other work in
other cohorts um with gestational
hypertension and other adverse pregnancy
outcomes I mentioned the xwas that we
did on cardiovascular disease and warned
you we were going to scale that up um so
we did um so we we took the same sort of
approach where again we're looking
across all of our surveys got that sort
of summarized here in a graphical
abstract and we ran xos over all of our
sort of common complex diseases that hit
a certain prevalence that made us feel
good about ourselves for for statistical
power um cut offs so you can see it's
the things that are high frequency
hypertension cholesterol um migraines
asthma type 2 diabetes Etc and we pushed
it through sort of what we now consider
our our xwas pipeline so we looked first
at single exposure modeling right
analogous to a gws I'm going to fit a
regression model for each one of my
exposures to each disease outcome then
we followed up with sort of
multi-exposure modeling so really just
sort of it's we use an algorithm called
DSA deletion substitution Edition it
it's sort of
nice nice um training testing um
validation framework for for sort of
last so
regression um and then think about sort
of our results visualization and
interpretation there here I've just got
a cartoon of forest plots with different
odds ratios um the other thing we wanted
to look at um in this study was also
looking at the correlations across our
um exposure um information right because
we know no one is exposed to one thing
at a time right back to that
exposomics um approach right exposomic
is everything's happening all the time
all at once always um so could we look
at sort of how these variables are
correlated and we didn't do anything all
that clever but we did something very
very tedious and a theme to my lab is we
do something tedious and probably take
that too far um but we wanted to look at
those correlations in survey data um we
didn't invent any new statistical
methods but we did the really rigorous
sort of hand cleaning that you need to
do for for survey data um so some survey
data is binary yes no you have this
disease you're exposed some of it is
ordinal um right where you've got you
know the equivalent of extra small to
large t-shirt sizes um in the data in
the way that you've collected it and if
you want to be rigorous about
correlations th those require different
things um so a survey question that you
have said are you exposed to asbestos
yes or no that's actually a binarization
of an underlying quantitative value and
so if you just take correlation types
for binary values you're not going to
get the right answers um so we did the
tedious thing you need to do if it
polychoric um correlations and um you
know like each one of those data type by
data type across those thousands of
exposures um so ran that correlation
analysis we also know different surveys
have different sample sizes um and a
correlations interpretation is is
heavily influenced by sample size so we
fit a a blup model to actually shrink
correlations so that we can put them um
on the same um scale there and and help
visualize those results here you're
seeing a um sort of a a spaghetti a plot
here right that will represent the
correlation across different
variables and what we really wanted to
do um was was make these results as as
available as as possible um so we ran
these xwas for all those different
outcomes and and we keep building on the
the diseases that we've evaluated and we
put all that data available in what we
call a pegs Explorer so if you ever
wanted to google ni and pegs it takes
you to our website where you can get
information on the study itself but also
go into some of our results explore um
so we've got an interactive web tool if
you're interested in this called pegs
Explorer and you can pull up for any of
the diseases and any of the um questions
that we have asked here you're seeing
example you can you can interactively go
through each one of those results see
what the association effect size was
we've got options for you to pick that
in different strata of the data with or
without different covariant um
adjustments and you can also download
we've got a visual toolkit that you can
look through um and you also can
download all that data um and and deal
with it computationally if you're
interested um so we keep building on
that like said then we've got these
correlation Globes um as we show them
that let you interact with how U people
are exposed to these these different
exposures um
jointly that's being a little weird with
the there we go and then the last um
study I'm going to sh talk about here in
pegs actually gets more towards the
Precision Health that I know this group
um is interested in so thank you for
your patience in some of the discovery
stuff um that we've done we'll talk
about um a recent study with a with a
clear Precision environmental health um
aspect and then what we're doing um in
that space right now that that's ongoing
work um so I've had an interest in in
diabetes have done a lot of
pharmacogenomics with diabetes drugs and
honestly the strongest signals we see in
pegs on any of those common complex um
traits tends to be
diabetes so here we wanted um to
directly compare sort of the the
performance of polygenic scores um which
have been the focus of human gene like a
lot of a focus of a lot of human
genetics work in the last decade um at
least right where what you're trying to
do is across the genome whether that is
a certain number of genes or really
you're trying to fit all of the genetic
variant across the genome and build a
predictive score um for your risk of a
certain disease um so we're interested
in comparing polygenic scores that have
gotten a whole lot of work with poly
exposure scores so hopefully you can
imagine what we mean by poly exposure
score so analogous to a polygenic risk
score we're going to look across those
exposures and build a risk score um that
sort of is associated and predicts your
disease outcome based on your
exposures so our first example here is
in type two diabetes um so um I'll I'll
I'll try to go through slightly
complicated plots um pretty quickly um
but so we compared a polygenic score um
there's there's a catalog of those
available I'll talk about that a little
bit more in a minute so there are some
well you know sort of well studied um
well used polygenic scores we fit our
own um as well to compare and then we
built poly exposure scores from our
survey questions um we also built a
clinical risk score for diabetes that
has been in the literature and this
clinical risk score in
risk score includes things like have you
ever been diagnosed with pre-diabetes
and BMI this is a very very strong
predictor um we did our xwas to help
explore some of those exposures and then
built a poly exposure score again with
really pretty simple sort of lasso based
um um methods with with training testing
we found some interesting things um in
this study with asbest and cold dust and
then sort of long story short what we
showed is the if here's my odds ratio um
of each of those scores so I've got my
clinical research score my polygenic
score then combinations of those my
environmental risk score long story
short the ex environmental risk score um
outperforms the polygenic score sarily
and even when I fit every one of those
different combinations the poly exposure
score always adds significant value
whether I want to look at that at AOC or
an NRI or whatever that fit um is so
even if I'm modeling with a clinical
score and a polygenic score adding that
um exposure score adds
value um our ongoing work and I'll show
you a little bit of hot off the presses
um data that came off literally this
week um to to present but we're
following up our xoss approach with the
growing number of GIS um exposures we're
interested in some genetic correlation
analysis um as we think about sort of
the association between variants and and
um exposures we're using that epigenetic
data for epigenome wide asso ation
studies um we're also following up with
epigenetic biomarkers so there there's a
lot recently out there on sort of Aging
clocks right your biological age versus
your chronological age um Based on
epigenetic data that that we're working
on I'm doing more um integrative omx
analysis and then what I'm going to talk
about here for a minute is the expansion
of that poly exposure and polygenic um
comparisons so as I mentioned early
earlier stole a graphic here for um
polygenic um risk scores right for
across your genome you can put sort of
um the the population you're interested
in and percentiles of risk right the
fifth percentile versus 99th um
percentile the same sort of
interpretation you would have for test
scores or or anything else um so so
thinking um in that way and we build the
poly exposure scores similarly so like
said this poly exposure score it's a
composite metric that Aggregates these
multiple types of exposures that that
you might see designed to capture the
broad and cumulative impact of diverse
potentially modifiable factors on health
or other outcomes and wants to emphasize
the inner playay of various exposures
rather than just one single right so so
really taking that exposomic that joint
um exposure and what does that do for
your
risk and in this study differently than
the diabetes one that we we published um
we've broken our our poly exposure
scores into two types so hang with me on
jargon because some of the jargon I'm
trying to invent um so um you can
imagine some of our exposures um are
things that that are well outside of any
individual participants control
place-based disparities things like
social determinance of health and there
we're we've built scores and we're
calling them poly social scores because
these really are aggregating social
determinants of Health that an
individual can't normally change you
can't tell somebody to move to a better
neighborhood they already would have if
they could have um right so these are
place-based disparities things like
income and socieconomic status housing
types and conditions social
vulnerability sorts of metrics and we
contrast that to what we're calling in
this study are poly exposure scores so
these are things that might include
things that really might be
Interventional at an individual level
those occupational residential hobby
based exposures that you could think
about buying new filters for your house
or mediating um some of that exposure so
think of them as more of these
Interventional factors at an individual
level occupational Hab hazards hobbies
and lifestyle choices and stress
exposures that are also
modifiable um so we started by
calculating polygenic scores so there is
a catalog there's a polygenic score
catalog that has a number of polygenic
scores that have been previously
published in other studies there are a
little over 3,000 of those that were
published um so we we downloaded those
and fit every one of those um to our
pegs participants so we've got these
multiple polygenic scores for each one
of our participants we looked at those
common complex traits that are again
above our prevalence cut off so at the
things that you saw for the exposome
wide Association um studies are now
being used in this and you can imagine
that these different traits have a
different number of polygenic scores
that have been published things like
type two diabetes that are more studied
had about 80 that we downloaded and
things like migraines only had a couple
of published um polygenic scores so keep
in mind there are a number of caveats to
that um but these were picked based on
the closest relevance to our phenotype
um and then we actually when I'm going
to show you results we're presenting the
results of whichever polygenic score
actually had the highest a in pegs so
we're trying to give these polygenic
scores a head like a head start right so
of those almost 80 that we fit for type
two diabetes I'm going to present the
results of the best fit um in pegs and
then compare that to our poly exposure
scores I'll go by this quickly if you're
familiar with an upset plot um it's
really easy to interpret and if you're
not it's a little hard to explain but
it's sort of a fancy Vin diagram um so
our our variables that came together for
the poly social score um versus the poly
exposure score we had a lot of overlap
and the same sort of social determinant
measures came up across diseases then
for the individual exposures there to
the right you see fewer connections um
down below across diseases that more
unique exposures um were pulled
in um overall again we did the same sort
of let's compare our polygenic scores to
these two poly exposure scores I'm going
to draw your attention to this graphic
here again on the X you've got the
traits that we've studied and on the Y
you've got the Au so the higher the dot
you see on this Dot Plot the stronger
the association the stronger um
potential prediction um and I've got
sort of the
different scores colored here um so the
the the purple are polyenic scores the
orange the poly social scores and the
green are the poly exposure scores long
story short what I want you to see here
is in every case the polygenic score has
much lower performance than either the
poly exposure or the poly social score
and there's really not a big difference
between the polyal and the poly exposure
score this becomes important when I talk
to Jan that's my copi and pegs she wants
to know not only will these exposures
predict the outcome but what's the
minimum number of questions I can ask
right are pegs participants answering
over a thousand questions isn't ever
going to be practical um for any amount
of translation but but can we get down
to sort of a minimal number of sets um
of questions which is which is
interesting I'll show you here just a
couple of other results if you're used
to seeing sort of an au or model fit in
a different way um here's our example
for type two diabetes again the purple
is the polygenic score and all these
others are our different poly exposure
scores and you can see in lower GI pups
you see even a stronger Gap um there
between the polygenic scores and the
poly
social I think it's pressing showing it
in one other way just to try to to drive
um the results home here you can look at
sort of prevalence plots so this
actually is a number of cases it's it's
um prevalence of that disease given the
different percentile scores for the
different scores so across a polygenic
score polyal and poly exposure what I
want you to see is that there's more
discrimination you get a higher range
from bottom to top there with the with
the poly social and poly exposure scores
then you do the polygenic
scores so I'll stop there I know that
was a whirlwind tour of things I've been
thinking about with pegs so thank you um
for your patience there I'm G to switch
gears now to talk about the the partner
research study um that Josh and all the
leadership at all of us um has been just
incredibly um supportive um and amazing
on so talk about that a little bit um so
stealing some slides from them I don't
even how I stole it from but it's
probably from you Josh um on sort of um
priorities in all of us of the different
um types of data that are they're
integrated right there's a wealth of
information in genomics and wearables
the BIOS specimens you all have
collected combined with a health record
different behavior and especially
fantastic social determinants of Health
surveys and then a number of physical
measurements how to sort of pull in more
and more environmental um data I I find
personally really exciting and doing it
for all all the reasons
um right sort of um pulling that sort of
data in helps how do socioeconomic
factors interact with exposures um for
risk um to really think about about you
know Health disparities think about risk
and prevention what environmental
biomarkers could identify risk for
future disease could these sort of
biomarkers be helpful in diagnostic what
sort of exposure signatures do you find
in patients with this or without this
different condition um treatments and
outcomes can we objectively obsess
whether treatment or intervention might
be effective right sort there's so many
broad goals you could think about when
you when you start thinking about
pulling environmental data into all of
us we had a fantastic Workshop um almost
three two and a half almost three years
ago right now where um all of us
leadership folks from NS and people from
the community at large um really spent a
few days thinking about sort of
priorities when you think about adding
in environmental data and a number of
themes um emerg sort of linking to
geospatial estimates linking to climate
and weather data exposomics whether you
mean that from surveys or mass spe based
sort of exposomics Technology thinking
about personal exposure risk and and
effective exposure on on diverse
populations and we followed that up um
with some really exciting um studies to
incorporate um the environment into all
of us um so really we sort of left that
with sort of a a three-phase um approach
to incorporating environment for phase
one and two we really based on
geospatial estimates and all of us has
made just an absolutely incredible
investment in getting addresses address
histories for participants and then
doing the geocoding um to link that data
um to exposures they're doing that
through a clad mechanism and we're
hopefully having some collaborations
with another niehs um effort that I'll
just mention um briefly and then phase
three is the partnered research study um
that I'm leading here we're here we did
a sort of a pilot case cohort study um
looking at incident diabetes um so
actually a ble to collect um sort of
untargeted mass spec based exposomics
and bios specimens from participants
that had incident diabetes versus a a
cohort um sample for comparison um which
is really exciting we're in the middle
of that data collection right now um I
had a meeting yesterday we done 51 out
of 71 um batches of collecting the
exposomics data so it's moving along
nicely oops I mentioned for the
geospatial estimates um there's another
effort at at
ni um chords is the acronym the the
acronym is under development again right
now um but it's it's a um Court F funded
um project to build tools to do
geospatial
linkages so it's taking data from NASA
from EPA from satellites um in imagery
and building tools to help people do
those linkages um both to do software to
build and curate web- based resources so
let's say you're a human gen IST that's
never done environmental Data before
been there um and you say maybe I'm
interested in air pollution it's
building web tools and AI tools to say
I'm interested in air pollution and it's
going to pull up what resources are
actually accessible to you here's the
EPA and how they do it here's Noah and
NASA and their versions right to to get
you started um on adding environmental
data into your
studies and I said specifically um this
all of us ancillary study um launched
last July we're working on identify we
identified participants um sent bios
specimens to to a laboratory here in
North Carolina that runs these Mass Spec
um sort of assays and that that data is
in progress now um so we'll have that on
the case cohort study it will go into um
the all of us data um Research Center
and their ecosystem so that data will be
available to everybody that has access
at the right tier within all of us which
is really
exciting I said it here I'm sort of
moving on phase three R said looking at
sort of those that exposomics and and
associations with outcomes you can
imagine we'll be doing things like
testing for associations of those
exposomic sparkers with development of
diabetes with the development of some of
the diabetic complications retinopathy
any of those others um with
comorbidities and coort alties um and
we'll have the wonderful opportunity to
um sort of look across MMS um this is
one of the largest exposomic studies
ever done um so just having this number
with this Geographic and demographic
diversity is pretty unprecedented um and
pretty exciting um so we're we're in the
middle of that right now the masspec
technology um that we've chosen it's the
here lab it's it's here um up the street
in Chapel Hill um collects a broad panel
of metabolites um whether you want to
call it metabolomics or exposomics so
we'll have a number of um sort of
endogenous um metabolites um
representing host metab ISM so amino
acids and polyamines and carnitine these
sort of host metabolites then we'll have
a lot that we can measure in the food um
metabone um caffeine and um all sorts of
um all sorts of different um sort of
subcategories of vitamins minerals Etc a
number of environmentally relevant
metabolites posos phols parens Bates
things you probably heard in in the news
if nowhere else and then additional
categories will'll be able to look at
drugs and medicines folate vitamins um
amino acids Etc this is what we're
expecting um off there so we can be
thinking about things that both use the
annotated um metabolites that we'll be
able to measure as well as thousands of
sort of Quantified metabolites that
don't have any annotation to them
yet and this can support a wide variety
of things so here's a very big slide
with right now and all of us we're going
to have genomic data exposomic data
we'll have Society data societal data
social determins of health and then
phenotypes so you can imagine um
hopefully some of the things I presented
that we've done in pegs I I hope some of
the similar thinking can come into all
of us where we wherever we're thinking
about xw whether that's from the
biochemical data or the geospatial data
can think about risk score analyses
again using all of these different
components um and to mention the the
long read sequencing has epigenetic
markers um for all of us too so can add
that um as well can think about all the
sorts of interaction analyses and then I
think a lot about variance
decompositions how much of this disease
is due to genetics and environment and
it's really exciting um like I said I've
had a whole lot of fun in pegs and it is
cute and adorable and being able to
think about some of these approaches and
something that has the mega sample size
of all of us um is just really really
exciting because that's that's what it's
going to take um to to look at Gene
environment
interactions all of this and more um
will be possible and I think that's
really
exciting and also I'm in it's
biology I'm suddenly echoing um I don't
know why but like said a lot of this is
going to need to be fueled by methods
development right there aren't tools out
there to ask and answer all these
questions um so there are a ton of
really exciting data science um
opportunities um as well that I I think
are important so I'll stop there um
again thank you for for dealing with
this Whirlwind um
I I I hope um it it was informative um
and sort of the ways we're thinking
about this of course I have to think all
sorts of people and there are more
people to thank um than just listed here
but of course have folks with the the
pegs leadership that I mentioned um I
see somebody already out there from our
external Advisory Board um even that I
that I see in the audience I said people
with the gis um all of our our
colleagues at all of us um that are just
amazing I said highlighting some of the
the computational work some of the
fromat assist that have done a lot of
this um and our our collaborators and
contract support and then first and
foremost the the pegs and all of us
participants that that give us generous
access to the data so I'll pause there
for questions and I'll
stop
sharing um there thank
you I don't know if you could hear that
but people were clapping for you you
could probably at least see them
clapping perhaps um so now it's open the
questions I have as as people you can
write them if you're online just write
them into the Q&A and if you are um here
you can just come up to the mic uh the
first uh I I'll ask first one that I got
online um and it was um how can we bring
this knowledge to the
clinics well I mean like we're we're in
sort of the discovery phase right the
the poly exposure score and things but I
I hope that will actually directly be
clinically applicable I I hope at some
point um and there some huge and
interesting efforts on adding social
determinance and other exposures to
electronic health records actually but I
I'm hoping this sort of um information
the sort of knowledge could help
clinicians um learn better who to screen
for what diseases or um you know which
questions to ask and how to take that
into more holistic um care so I know
we're a long way from that um but but I
hope we're less far away from
that than than further if if that
helps yeah it makes me think of the uh
all the qu sections of our notes and and
talking about social history how we
could really expand that as well another
one online um is uh do you have
sufficient data for the vast numbers of
variables to use deep learning methods
to enable more agnostic
results yes do we have the sample size
right anytime you ask a statistician if
they should have more samples we will
say yes it's sort of like asking a dog
if they're hungry like yeah the dog's
hungry and I would love more samples um
but we absolutely are able to use some
of those those new machine Learners um
and I've got some folks working on that
I always think of um right anytime
you're doing hypothesis generation I
want to know what would be a useful
hypothesis to follow up um so met like
my dissertation 20 years ago involved
neural Nets um I've been doing that sort
of machine learning for a really long
time um so you have to think there are
cases where that sort of prediction that
sort of multi-omic integration is is is
really useful um and could could be sort
of predictive power um so we've got some
folks working in that space with that
sort of machine learner and then I
sometimes contrast that to the stuff
that that maybe I was more heavy on here
where we're trying to do hypothesis
generation that we could we could follow
up um right back to the clinical
question we are long way away from
somebody spending tens of thousands of
dollars for multiomic in a clinical um
setting so so we do we have one Avenue
that I haven't talked about a ton with
with machine learning and with methods
to narrow down the search base for G by
followup that's more targeted that that
relies heavily on machine learning um
but yeah we we can wield those tools um
in this this sample size for sure
great hi um thank you so much for your
talk I have a maybe not completely
formed question but um in looking at all
of the awesome survey data that you guys
have I think I first thought about how
you get people to answer those surveys
um maybe particularly people who live in
areas that don't have great internet
access or things like that because um I
went to college in North Carolina and
something we talked about a lot was the
poor health outcomes in the western part
of the state so just wanted to see if
you could touch on that and maybe talk
about any challenges you've had with
like follow up with participants from
those areas and things like that
absolutely yeah it's absolutely a
challenge so a lot of the recruitment
was done with at at the beginning um
when pegs was starting was done with
very targeted recruitment drives um and
so that involved things like
Partnerships with church churches in
those like rural areas that are a
community that is trusted and that could
interact um there were also a lot of
joint recruitments in other initiatives
um a lot of like women's health sort of
initiatives so some of it we have more
women um because women tend to volunteer
for things but we also had some
recruitments tied into some Women's
Health um initiatives we had massive
challenges during covid like a lot of
people of people moving and relocating
we do a lot of followup electronically
now um right like we can take your
addresses and like every adverti like
there are plenty of you know
computational resources that track who's
where and and can reach back so we
definitely had a gap in covid with
people moving um and sort of reaching
back out um as well for the most part um
I'll make I'm gonna sound really
flippant but I'm not these participants
will just volunteer for anything like
our participants are just amazing I
don't know like I almost worry about
them sometimes like you sure okay you'll
keep going so also now we've switched a
lot to electronic you know like we're
entired electronic surveys um in sort of
our ongoing followup um sort of
callbacks um where we incentivized
people to take the surveys whether it's
with Amazon gift cards or or other sorts
of of sort of simple things the other
thing I will highlight about pegs that I
haven't mentioned and just use this as
an excuse is all of these participants
are available to call back in our
clinical Center um at niehs so pegs has
also served as almost a core facility um
for people that need uh bios specimens
or or specific things we've had people
do studies where we gotten we've called
pegs participants back in to do skin
biopsies on sun exposed versus unexposed
regions and looked at mutation rates
we've had participants call back for
inhalation assays like they'll sit in a
body pod and breee breathe diesel
because we ask them um nicely so it's
it's a really
engaged community that cares a lot about
it so in some ways I'm like I don't know
the secret outside of like these people
are just really really generous with
their their time and
energy awesome thank you
another one
online hi Allison great talk does pegs
have disease incidence data are only
prevalence are there efforts to bring in
longitudinal data to pegs yes right now
we have prevalence um it has been an
ongoing effort but we're finally it's
maturing we are linking to health
records um first with UNCC Chapel Hill
um so once we're there um we've we've
done a lot of the the data sharing we've
had a lot of IRB and reconsent um to get
this done but I'm hoping in 202
will have electronic health records um
and UNCC Chapel Hill doesn't just have
their health system isn't just the
hospital but like a large number of
primary care um places are are included
in that UNCC so it'll be it'll suddenly
have longitudinal data sort of in the
same way that all of us does um as we
work to finalize that that health record
so I'm very excited for
that great great talk Allison bless
beicker um question for you on how this
really goes forward my you know head
explodes with a number of variables
you're talking about here and you know
in of course in
genetics it's taken us 15 to 20 years to
go from GW
was associations to people actually
starting to work out cause and effect of
these variants and it's been a really
hard
slog and I can't even begin to imagine
what it that equivalent looks like for
you from this poly
exosome score to get to cause and effect
or how what it is in that score that is
disrupting the physiology that makes
these people sick what what are you
thinking about what that looks like yeah
you know
I understand and have
no no easy solutions there right um and
I I know you don't think we do um yeah
like I I think some of the issues there
there's some things that are way easier
in genetics right like I know how to
format that like I've got nice
structured data it took us a while in
the gws era for chips to evolve and get
to you like there were challenges but I
I can harmonize genetic data from the UK
from Canada from all of us trivially by
comparison of harmonizing environmental
data um as well so there are all sorts
of issues with sort of um data sharing
in in exposomic and in Environmental
Health Sciences um I I said we that pegs
explore some of those things like for
the most part if you want an
epidemiological study there's so many
data use committees and people have
already like pinned what their little
niche is um for this area so I think I
think data sharing is going to have to
be more expected for us to get there um
a more open fisted approach that I think
genetics has had more than Environmental
Health Sciences and I don't triv
trivialize that there are different
issues to consider with that I'm not
trivializing that that challenge um but
I think that needs to happen um I think
the opportunities for people in
leadership at all of us in UK bio bank
and other International efforts like a
willingness to try to standardize at
that level would go very very far um for
that that like G by e or E is
notoriously hard to find right for all
the statistical challenges and all the
high-dimensional challenges and honestly
the mega sample size is what it's going
to have to take to get there of just
these huge mega samples within clever
people doing causal inference in
observational studies to help narrow
that down and always really clever
Partnerships between basic
biologist mechanistic biologist
toxicologist what yeah like it it's just
going to be even more interdisciplinary
to to get somewhere
meaningful is is that helpful or sort of
in line with what you're um it's very
analogous to how we were what we were
saying when jwa started like my
God
I I'm gonna take a prerogative to ask a
follow-up question uh based on that that
i' had written down which was it's
interesting how similar the poly social
and the you know poly exposure scores
were in predicted value one being
intervenable one not I was wondering if
that's actually like just similar
predictive
values or it's actually the same similar
for the same person and if so what does
that actually mean does that mean that
we tend to Cluster our exposures in some
way you know that are the ones we
intervene on not or you know is there
something we can learn from that
basically yeah we're just exploring this
like some of this is hot off the presses
um I would say it's more surprisingly
the same predictive power than
necessarily the same individuals um they
are certainly correlated AB absolutely
there is correlation but it is the
predictive powers are more similar than
I would have expected when I just looked
at the correlations um but we we're
really just starting to pick all of that
um we're we're starting to disentangle
this not just for diabetes but more more
broadly yeah the predictive performance
is stronger than I would expect on the
correlation Neil I think you get the
last question oh this one will be um
relatively simple so if I heard you
right you said that you were in 99 of
the 100
counties what what is it about the onean
county that you're not
in I'm not going to remember the name of
it I'm from North Carolina so I should
it's actually the one to the furthest
Northern right that is just really rural
um it's like right at the edge of the
Outer Banks there are not many people
there and we want to recruit in that
county for you so um but if I need to go
to the beach if I need to go to the
beach to to really make sure we're there
I would I would make that sacrifice I
would I would go to the beautiful
beaches if I had to Great uh thanks for
all the folks that ask questions online
and it's also um it's been a great talk
Allison I really appreciate uh you
coming and giving this I think you give
a a exposure to a kind of thing we don't
get to talk a lot about um and the
importance uh we often talk the
importance of zip code and code but this
is like a whole different layer on that
as we talk about lifestyle environment
and biology here so thank you very much
for joining us thank you this is lovely
and we're clapping one more time if you
can't tell thank you thank you