Submind YouTube summaries
Thumbnail for CPHR Seminar Series - Alison Motsinger-Reif

CPHR Seminar Series - Alison Motsinger-Reif

Watch on YouTube

Video summary

Alison Motsinger-Reif, Chief of the Bioinformatics and Computational Biology Branch at NIEHS, presented an in-depth overview of the Personalized Environment and Genes (PEGS) study and its integration with the All of Us ancillary research. The PEGS cohort comprises approximately 20,000 individuals in North Carolina who have undergone genome sequencing and provided extensive data through three massive surveys containing over 1,700 questions covering clinical conditions, lifestyle factors, diet, sleep, and residential or occupational histories. By utilizing geospatial linkages, the study maps participant addresses against environmental hazards, air pollution, land use, and climate data across 99 of North Carolina's counties, creating a robust foundation for analyzing the complex interplay between genetics and the environment. The presentation highlighted several key research themes, including Exposome-Wide Association Studies (XWAS) that identified significant links between blood type A Rh-negative and heart attacks, as well as paternal education levels and cardiovascular outcomes. Researchers also found strong effects of air pollution mixtures on autoimmune skin diseases comparable to smoking history, while population-specific analyses revealed different HLA gene classes associated with asthma in African American versus European ancestry participants. Furthermore, despite small sample sizes, genome-wide significant hits were identified for rare outcomes like gestational hypertension, and poly exposure scores aggregating environmental factors demonstrated superior predictive power for Type 2 diabetes compared to traditional polygenic risk scores, even when combined with clinical and genetic data. To support these findings, the team developed "PEGS Explorer," an interactive web tool that allows users to explore associations, stratify by covariates, and download data, while the study is transitioning from cross-sectional prevalence data to longitudinal insights through linkages with electronic health records. The scope of measurable metabolomic data includes host metabolites, dietary components, environmental pollutants, and pharmaceuticals, which are integrated with genomic, epigenetic, and social determinants of health to conduct risk score analyses. Addressing recruitment challenges in rural areas, the study successfully partnered with trusted community institutions like churches and utilized women's health initiatives, electronic follow-ups, and incentives, while serving as a core facility for clinical procedures such as skin biopsies and inhalation assays. Looking toward the future, the All of Us ancillary study aims to integrate deep exposomics via mass spectrometry and biospecimens to investigate incident diabetes and complications, leveraging upcoming collaborations with NIH Common Fund projects. Motsinger-Reif emphasized that while moving from association to causality in exposomics is more difficult than in genetics due to the complexity of harmonizing environmental data, current efforts focus on hypothesis generation and screening strategies rather than immediate clinical application. She called for greater openness in data sharing and interdisciplinary collaboration involving basic biologists, mechanistic biologists, and toxicologists to further disentangle correlations between poly-exposure scores and poly-social scores, ultimately advancing the understanding of how environmental factors interact with human biology.
Read the full video transcript
welcome everyone I appreciate everyone that's joined us today both uh in person and online uh today we have a outside uh guest to speak but from the NIH I'm from our North Carolina Branch at National Institute for Environmental Health Sciences aliser Allison moninger rif will be joining us and presenting her work um from nihs and and uh she's done a lot of work at the intersection of genomics environment and a number of different variables she is c-i of a cohort called pegs which I should have actually looked what this stood for but she probably going to mention it and it's an environmental uh collection of about 20,000 people in the North Carolina area that also has genome sequence data it has a deep environmental exposure data and she's looked at that data in a number of different ways that she'll talk about she's also leading a um a study that is an anary study to all of us which he'll talk about which is gathering uh deep uh deep um expose them data to link with all the other types of data we have to look at the risk of incident type two diabetes in the environment she is a let's see your uh their chief of your branch which is the uh bio statistics and computational biology Branch at n niehs and been a great collaborator in these projects so Allison I'll turn it over to you thank you so much thank you Josh for the invitation um thank you everybody for their time especially as many things are going on right now I nobody's everybody's time and attention is even scarcer than normal so I'm really grateful um to get this time with you guys um I'm going to attempt to share my screen now and hopefully we'll make that transition just fine right do you guys see one big slide excellent love it when the magic works um so as Josh said I'm a human geneticist um at nihs um I come from a human genetics background um Josh and I are both vendy alums um at different times so I think I'm supposed to say anchor down to him um but that that slang uh is after when I actually left um and so it's it's been really fun like I said coming from a human genetics perspective I was a faculty member at NC State for a long time in their statistics department and B informatics Research Center and then I joined in IHS about 6 years ago um right now for the leadership role in the branch and have had the wonderful opportunity um to work with this study um that Josh mentioned pegs stands for the personalized environment and genes study and I'll tell you a little bit more um about it as we go along so overall um I've gotten away with in my career sort of a mix of methods developments um from a statistical biostatistical computational biology um point of view as well as really doing a lot of Applied work in the the the balance of my portfolio has changed Through The Years um but really enjoying doing both of those things um I'm not that clever a statistician I'm not going to be in my office deriving the asymptotics to anything new I steal all my methods um challenges from the real data um that I'm working on so I'm always really transparent and I guess if you highlight some themes I've worked in machine learning pharmacogenomics thinking about Gene Gene and Gene environment interactions dose response modeling and really I'll deal with whatever m is huge and popular and um and challenging so so this data I'm going to talk about now has been a real fun way to to really think about diverse data um I hope it's okay um I'm going to take you a little bit on a whirlwind tour um here and appreciate your patience I'm going to tell you some of the things we've been doing over the last two to three years um with this study that I've mentioned um from a number of different angles um sort of the the studies I'm going to talk about are prioritized based on mostly fellows and trainees interest and where they want to go in their career as well as what we think um is going to be scientifically um impactful and as you mentioned I'm going to talk first about the work we've been doing in pegs and then we'll end the discussion with some of the work that's ongoing in the the partner research study um that Josh mentioned with all of us and I hope you'll be able to see the connection between some of the kinds of work and the way we've been thinking about things um in the pegs data and how that um Can can be really applicable in the types of data that that all of us is now really thinking about about collecting um that I think is is really exciting so I hope you're you're willing to hang with me um we'll we'll move around um across topics um so I'm going to use the the phrase we um very generously right like any good Pi I actually don't do anything um everything that I present as we um is from a really talented um group of folks um that have done the the bulk of this work Dr Freda OCT is a um contractor staff scientist Dylan Lloyd is a grad student that just graduated and we'll be moving to Harvard next month for a postto John house is a a staff scientist Jasmine Mack's another graduate student that'll defend this summer and then Ryan Campbell is another um contract staff member and Joe is a is a postto so really I'm going to take credit for all their work as well as I have to go on and highlight other key co-conspirators um that help me get away um with all of this fun stuff um the folks collaborating in pegs Jan Hall has been the co-pi of pegs so she's retiring next month Dave Fargo Charles Smith and Messier both are are all fantastic bioinformaticist computational scientists um tenure track investigators then I'll highlight Jeff and Josh and others at all of us um that have really been the the co-conspirators on the ancillary study Dr Rick week our Institute director and one-on-one collaborator um at niehs that has supported all this work um as well so I'm just going to sort of introduce here I don't I don't know what my audience is on the mix of geneticists versus other interdisciplinary scientists but I want to sort of cue up what we're doing within sort of an exposomics framework um so like I said I was trained initially as a geneticist so right I'm certainly um used to thinking about what are the what's the genetic ideology of common complex traits um but as I did more pharmacogenomics and moving more and more into Environmental Health Sciences um really am appreciating how much more information there is to gather from an exposomics um standpoint and so when I say exposomics I'm going to be really really inclusive in that so here I mean things like that you probably think about as environmental exposures what's your physical environment how much air pollution is near your house or near your work what sort of chemicals um do you use in your job occupationally every day so so really broadly but I also want to make sure we include in that the social environment whether it's social determinance of Health um or other um sorts of of social factors um that that control the the body's physiological response um things that you might consider lifestyle diet exercise sleep that kind of thing and then also within exposomics the the body's response so how do does the body manage that whether that's exposomics metabolomics right what how does the body process um these chemicals that they're exposed to um Etc and hopefully you'll see I I'm very inclusive in that and sort of an exposomics approach exposomics data to me and the way I'm going to use it can come from multiple sources so that could be things like survey questions that I'm going to talk a lot about or that could come from geospatial estimates of exposure right how do we we've got estimates of air pollution or other sorts of exposures that could um include you know sort of omix data right metabolomics exposomics um as a as a as a proxy for your environmental exposure so try to think about that that really inclusively so now hopefully you're sort of queuing up to the framework and I'll give you some more details on the peg study um and and talked about what we're doing with it so as Josh mentioned this is It's a relatively small cohort especially from a genomics um perspective but it's a decently large cohort from an environmental health sciences perspective and we've spent the last five or six years really Gathering these multiple data types um in the study so we've got a lot of of data on genetics um over the years this cohort has been around for almost 20 years and there's been some halfhazard sort of candidate Gene collection along the way um but a major effort a few years ago to collect hold genome sequencing data on everybody we could afford at the time which is about 4700 individuals we recently last year we able to get um methylation data epigenetic data from the Epic chip on everybody that we have whole genome sequencing um for phenotypes um we have a number of self-reported diseases and conditions um from our questionnaires and I'm going to tag in more detail on our questionnaires when I talk about the environmental exposure so pegs has huge questionnaires patently absurdly huge questionnaires compared to to what's often collected so we have three questionnaires one we call Health and exposure um and this gets to a lot of has a doctor ever diagnosed you with a b or c and then in your work or in your home have you ever been exposed to X Y or z a lot of questions stolen from well-established instruments studies like inhes nurses Health Etc and then we have two other surveys that were taken by fewer individuals but really extensive on what we call our internal and external exposome surveys so the internal exposome you could probably relabel lifestyle and that fit pretty well so what medications are you taking what physical activity do you participate in um that's sleep um stress that sort of thing then the external exposome has a lot of detailed residential and occupational um exposures and then we also have a continuously growing um and really dynamically growing right now um set of Expos meur es from geospatial linkages so we have addresses and address histories for all of our participants in pegs including things like the longest lived childhood and the longest lived adult address where we've linked a growing number of exposure estimates just to give you a snapshot of some of the demographics of pegs the average age when they took the health and exposure um survey is about 50 we've really got some pretty broad diversity in terms of Education levels you can see some graphs um there um we are mostly a Caucasian um ancestry cohort about 60% but almost 28% come from africanamerican ancestry which is which is decently unique actually um we've got more women than we do men it's a volunteer cohort that happens um pretty steadily and then we do actually have a pretty good diverse range of sort of incomes um and other um sort of measures of socioeconomic status just to give you a couple snapshots on those different surveys um so in that health and exposure um survey we asked questions about over 120 sort of clinical conditions and disease States we've got some snapshots here um you can imagine we have sort of about population prevalence um of most of the common complex diseases were slightly healthier than the North Carolina population um but but not by much um and like said to give you um just a snapshot of some of those those lifestyle um factors as well then I mentioned these exposome surveys are external and internal exposome um over 200 questions each exposome has things like the characteristics of your current and past homes workplace characteristics chemical and metalic exposures hobby exposures nuv light exposures and then internal exposome has medications vam vitamins supplements uh drug treatments um physical activities stress infection sleep diet um s sibling twin birth order um genetic history in total between those three surveys our our participants actually answer about 1,700 questions um which is anybody's everever been involved in sort of that that design that's that's a lot so it's really really rich survey questions then I mentioned our GIS data we've got a growing number of data layers there our first linkages were from to what I've learned is actually a technical term environmental bads so it's how far away do you live from these environmental bads things like airports Kos are caged animal feeding operations that are not entirely unique but a North Carolina specific um exposure and then things like cell towers drinking water dry cleaners hazardous wa sites um you can see the list um here I've just said distance to these um environmental bads um for again all of our addresses now and longest lived um adult and then longest lived childhood if that data is available the gis data we've got a growing number of linkages so we've done a lot of linkages to measures of neighborhood socio economics and structure um a lot to air pollution whether that's wildfires or pm2.5 again a growing number of those point source metrics distance to environmental bads um we're working on more and more particulate matter linkages um a lot of high resolution land use and land covers um measurements things like green space Urban density um a number of climate and climate change um measures including temperature extremes and and seasonal um variability and and a lot of sort of linking with with public open data API sorts of things we have great coverage over the the entire state of North Carolina we definitely are biased to to sort of the Research Triangle Park area here where NS is located but we have participants in 99 out of the 100 counties in the state of North Carolina I also am from North Carolina this is where I grew up um so it has lots of sort of personal resonance um for me now I'm going to start this Whirlwind tour of some of the things um that we have worked on and thought about um so I I hope this you know helps show the the kinds of things we're we're thinking about so former proo Dr Unice Lee um LED some work she's really passionate about cardiopulmonary um outcomes and wants to do that um for her career and led what we call an exos on sort of atherogenic cardiovascular disease outcomes so exas meaning exposome wide Association study you can think of this as analogous to a Goos if you're familiar with that area but we're really in a sort of a hypothesis Generation Um and sort of Big Data exploration space so she took all of our exposures from the questionnaires that over 1700 and did Association analysis where she tested you know exposure by exposure with association with one of several atherogenic cardiovascular um outcomes I'm showing you here a tile plot um so on the the bottom x axis you see the the exposures that that question on the Y AIS you see the cardiovascular outcome that's associated with and everything that has uh a colored thing on this tiop plot had a significant Association after a pretty rigid FDR multiple testing control the top actually has the effect sizes um of those and the bottom just has it magnified by by the strength of Association um so doing this um found some some interesting results we were able to reproduce a lot of things that make us feel sort of solid about the cohort and the results we're finding um and found some new things that hadn't been reported um before in particular strong associations with blood type um a Rh negative so a negative with heart attack paint related exposures with stroke biohazardous materials um at work with arhythmia and strong associations with paternal education level and several of those those outcomes so you can start to see we can start explore what are shared environmental associations and what are unique um and sort of look across that space you'll see that we've scaled up this xwat approach to multiple outcomes um following her work we've also done some very targeted sort of Gene environment analysis work um so here we we worked with canidate genes so we were interested in what's the association of the sort of distance to these Kos we call them these caged animal feeding operations which are really really they're they're what they sound like they're they're very dense industrial agricultural um sites that have absolutely terrible um environmental exposures there's a a town to the east um here called Smithfield that when you drive through it you smell the exposure for the Kayo um that is there your tires will smell like that Kayo even when you get home to the garage and Raleigh you have to spray them off if you drove through Smithfield um so these really are are pretty dramatic sort of exposures so we're interested in associations with IM immune mediated diseases and that distance um to these caged animal feeding operations in particular looking at um Gene environment interactions with known genes that are associated to these these disorders so we looked at individual diseases as well as grouped diseases um and what I'm showing you here is a pretty dramatic example of a gene environment interaction so on the y- AIS here you've got the distance to the Kos so this is how far away our participants live from one of those um kfo facilities our y has the odds ratio so you're seeing our effect size um and here I'm showing you the the sort of what the environmental effect is um alone and then I'm showing you that environmental effect in odds ratio by genotype for individuals with this um the homozygous minor alal you see a very very strong um interaction effect um so we're interested in report these on multiple sort of genes in ptpn22 um ahr pathway Gene interactions with with Kos um exposures that were really um convincing so like at G by E we're also interested in those geospatial um linkages that I mentioned to you and uh Kyle Messier that I mentioned there is really an expert in developing the the actual estimates um of exposure um so right here we're looking at air pollution you can imagine there this this data comes from different EPA monitoring sites and there's not an even distribution of of those monitors so we have to fit and um estimate these different components of air pollution and you can see a bunch of North Carolina maps um with the values um for those different components of um air pollution and he takes really sophisticated beian modeling sort of approaches to associate those air pollution um measures with um sort of in this case the the traine was really interested in autoimmune skin diseases and long story short found some very strong associations when you looked at the mixtures so if you looked at these individual components you didn't see the effect but when you're able to jointly model um the different components of small um particulate matter um air pollution was was able to see this and to to give you a little bit of an anchor um I said I know I'm going through it quickly but sort of the strength of these associations were about the same magnitude of whether or not you had a history as a smoker which comes up in autoimmune diseases and many other um common complex traits um all the time this really is a strong magnitude um of effect like said on par with whether or not you smoke additionally we mentioned we've got whole genome sequencing um and so we've been able to call the hlaa MHC region um to six-digit resolution and again this is work from Unice Lee the the woman that was doing the xos studies and did MHC region mapping with adult onset um asthma so she tested population specific specific snip and HLA alal and then protein um associations and did it um stratified for the two different genetic ancestry groups um that we've got here um and long story short she found some really interesting sort of population specific effects um so really found that in um sort of our participants with um African-American ancestry is really H and HLA do o I'm I'm sorry reverse that um in our European subjects it was H and H O and that class one HLA La genes were strongly associated with our participants of the European ancestry and it was actually class two genes for participants with African ancestry um which is interesting and and we're we're building on a lot of sort of the the pipelines and work um she developed for us for that HLA MHC analyses for for stuff that we interested in in all of us cohort um as well um another really talented graduate student um wanted to to pursue um gwas analysis with adverse pregnancy outcomes like I said I'm a human geneticist I was incredibly skeptical about sort of the the power um of this study to do any gws mapping and especially in sort of rare outcomes like adverse pregnancy outcomes but it's one of those moments I'm glad I didn't listen to me and I listened um to her which happens to me all the time um but she was really interested in that um so did GS mapping in pegs of gestational hypertension you can see we've got a quite a small um sample size um when you consider sort of what you're used to for gwos mapping um but here you can see that Manhattan plot down to the left we've got a genome wide significant hit and a couple of genes and including this um rare B Gene it's a retinoic acid signaling Gene um we were able to replicate that in the UK um biobank you can see there um sort of the the lead snip there um and really interesting rare be makes a ton of biological sense so it's um been studied and shown to be Associated um with sort of severity of protura um and and in preeclamptic studies um particularly it come up in the fetal genome but we're seeing it here now in the maternal genome um so that paper was recently published um and she's going to continue um other work in other cohorts um with gestational hypertension and other adverse pregnancy outcomes I mentioned the xwas that we did on cardiovascular disease and warned you we were going to scale that up um so we did um so we we took the same sort of approach where again we're looking across all of our surveys got that sort of summarized here in a graphical abstract and we ran xos over all of our sort of common complex diseases that hit a certain prevalence that made us feel good about ourselves for for statistical power um cut offs so you can see it's the things that are high frequency hypertension cholesterol um migraines asthma type 2 diabetes Etc and we pushed it through sort of what we now consider our our xwas pipeline so we looked first at single exposure modeling right analogous to a gws I'm going to fit a regression model for each one of my exposures to each disease outcome then we followed up with sort of multi-exposure modeling so really just sort of it's we use an algorithm called DSA deletion substitution Edition it it's sort of nice nice um training testing um validation framework for for sort of last so regression um and then think about sort of our results visualization and interpretation there here I've just got a cartoon of forest plots with different odds ratios um the other thing we wanted to look at um in this study was also looking at the correlations across our um exposure um information right because we know no one is exposed to one thing at a time right back to that exposomics um approach right exposomic is everything's happening all the time all at once always um so could we look at sort of how these variables are correlated and we didn't do anything all that clever but we did something very very tedious and a theme to my lab is we do something tedious and probably take that too far um but we wanted to look at those correlations in survey data um we didn't invent any new statistical methods but we did the really rigorous sort of hand cleaning that you need to do for for survey data um so some survey data is binary yes no you have this disease you're exposed some of it is ordinal um right where you've got you know the equivalent of extra small to large t-shirt sizes um in the data in the way that you've collected it and if you want to be rigorous about correlations th those require different things um so a survey question that you have said are you exposed to asbestos yes or no that's actually a binarization of an underlying quantitative value and so if you just take correlation types for binary values you're not going to get the right answers um so we did the tedious thing you need to do if it polychoric um correlations and um you know like each one of those data type by data type across those thousands of exposures um so ran that correlation analysis we also know different surveys have different sample sizes um and a correlations interpretation is is heavily influenced by sample size so we fit a a blup model to actually shrink correlations so that we can put them um on the same um scale there and and help visualize those results here you're seeing a um sort of a a spaghetti a plot here right that will represent the correlation across different variables and what we really wanted to do um was was make these results as as available as as possible um so we ran these xwas for all those different outcomes and and we keep building on the the diseases that we've evaluated and we put all that data available in what we call a pegs Explorer so if you ever wanted to google ni and pegs it takes you to our website where you can get information on the study itself but also go into some of our results explore um so we've got an interactive web tool if you're interested in this called pegs Explorer and you can pull up for any of the diseases and any of the um questions that we have asked here you're seeing example you can you can interactively go through each one of those results see what the association effect size was we've got options for you to pick that in different strata of the data with or without different covariant um adjustments and you can also download we've got a visual toolkit that you can look through um and you also can download all that data um and and deal with it computationally if you're interested um so we keep building on that like said then we've got these correlation Globes um as we show them that let you interact with how U people are exposed to these these different exposures um jointly that's being a little weird with the there we go and then the last um study I'm going to sh talk about here in pegs actually gets more towards the Precision Health that I know this group um is interested in so thank you for your patience in some of the discovery stuff um that we've done we'll talk about um a recent study with a with a clear Precision environmental health um aspect and then what we're doing um in that space right now that that's ongoing work um so I've had an interest in in diabetes have done a lot of pharmacogenomics with diabetes drugs and honestly the strongest signals we see in pegs on any of those common complex um traits tends to be diabetes so here we wanted um to directly compare sort of the the performance of polygenic scores um which have been the focus of human gene like a lot of a focus of a lot of human genetics work in the last decade um at least right where what you're trying to do is across the genome whether that is a certain number of genes or really you're trying to fit all of the genetic variant across the genome and build a predictive score um for your risk of a certain disease um so we're interested in comparing polygenic scores that have gotten a whole lot of work with poly exposure scores so hopefully you can imagine what we mean by poly exposure score so analogous to a polygenic risk score we're going to look across those exposures and build a risk score um that sort of is associated and predicts your disease outcome based on your exposures so our first example here is in type two diabetes um so um I'll I'll I'll try to go through slightly complicated plots um pretty quickly um but so we compared a polygenic score um there's there's a catalog of those available I'll talk about that a little bit more in a minute so there are some well you know sort of well studied um well used polygenic scores we fit our own um as well to compare and then we built poly exposure scores from our survey questions um we also built a clinical risk score for diabetes that has been in the literature and this clinical risk score in risk score includes things like have you ever been diagnosed with pre-diabetes and BMI this is a very very strong predictor um we did our xwas to help explore some of those exposures and then built a poly exposure score again with really pretty simple sort of lasso based um um methods with with training testing we found some interesting things um in this study with asbest and cold dust and then sort of long story short what we showed is the if here's my odds ratio um of each of those scores so I've got my clinical research score my polygenic score then combinations of those my environmental risk score long story short the ex environmental risk score um outperforms the polygenic score sarily and even when I fit every one of those different combinations the poly exposure score always adds significant value whether I want to look at that at AOC or an NRI or whatever that fit um is so even if I'm modeling with a clinical score and a polygenic score adding that um exposure score adds value um our ongoing work and I'll show you a little bit of hot off the presses um data that came off literally this week um to to present but we're following up our xoss approach with the growing number of GIS um exposures we're interested in some genetic correlation analysis um as we think about sort of the association between variants and and um exposures we're using that epigenetic data for epigenome wide asso ation studies um we're also following up with epigenetic biomarkers so there there's a lot recently out there on sort of Aging clocks right your biological age versus your chronological age um Based on epigenetic data that that we're working on I'm doing more um integrative omx analysis and then what I'm going to talk about here for a minute is the expansion of that poly exposure and polygenic um comparisons so as I mentioned early earlier stole a graphic here for um polygenic um risk scores right for across your genome you can put sort of um the the population you're interested in and percentiles of risk right the fifth percentile versus 99th um percentile the same sort of interpretation you would have for test scores or or anything else um so so thinking um in that way and we build the poly exposure scores similarly so like said this poly exposure score it's a composite metric that Aggregates these multiple types of exposures that that you might see designed to capture the broad and cumulative impact of diverse potentially modifiable factors on health or other outcomes and wants to emphasize the inner playay of various exposures rather than just one single right so so really taking that exposomic that joint um exposure and what does that do for your risk and in this study differently than the diabetes one that we we published um we've broken our our poly exposure scores into two types so hang with me on jargon because some of the jargon I'm trying to invent um so um you can imagine some of our exposures um are things that that are well outside of any individual participants control place-based disparities things like social determinance of health and there we're we've built scores and we're calling them poly social scores because these really are aggregating social determinants of Health that an individual can't normally change you can't tell somebody to move to a better neighborhood they already would have if they could have um right so these are place-based disparities things like income and socieconomic status housing types and conditions social vulnerability sorts of metrics and we contrast that to what we're calling in this study are poly exposure scores so these are things that might include things that really might be Interventional at an individual level those occupational residential hobby based exposures that you could think about buying new filters for your house or mediating um some of that exposure so think of them as more of these Interventional factors at an individual level occupational Hab hazards hobbies and lifestyle choices and stress exposures that are also modifiable um so we started by calculating polygenic scores so there is a catalog there's a polygenic score catalog that has a number of polygenic scores that have been previously published in other studies there are a little over 3,000 of those that were published um so we we downloaded those and fit every one of those um to our pegs participants so we've got these multiple polygenic scores for each one of our participants we looked at those common complex traits that are again above our prevalence cut off so at the things that you saw for the exposome wide Association um studies are now being used in this and you can imagine that these different traits have a different number of polygenic scores that have been published things like type two diabetes that are more studied had about 80 that we downloaded and things like migraines only had a couple of published um polygenic scores so keep in mind there are a number of caveats to that um but these were picked based on the closest relevance to our phenotype um and then we actually when I'm going to show you results we're presenting the results of whichever polygenic score actually had the highest a in pegs so we're trying to give these polygenic scores a head like a head start right so of those almost 80 that we fit for type two diabetes I'm going to present the results of the best fit um in pegs and then compare that to our poly exposure scores I'll go by this quickly if you're familiar with an upset plot um it's really easy to interpret and if you're not it's a little hard to explain but it's sort of a fancy Vin diagram um so our our variables that came together for the poly social score um versus the poly exposure score we had a lot of overlap and the same sort of social determinant measures came up across diseases then for the individual exposures there to the right you see fewer connections um down below across diseases that more unique exposures um were pulled in um overall again we did the same sort of let's compare our polygenic scores to these two poly exposure scores I'm going to draw your attention to this graphic here again on the X you've got the traits that we've studied and on the Y you've got the Au so the higher the dot you see on this Dot Plot the stronger the association the stronger um potential prediction um and I've got sort of the different scores colored here um so the the the purple are polyenic scores the orange the poly social scores and the green are the poly exposure scores long story short what I want you to see here is in every case the polygenic score has much lower performance than either the poly exposure or the poly social score and there's really not a big difference between the polyal and the poly exposure score this becomes important when I talk to Jan that's my copi and pegs she wants to know not only will these exposures predict the outcome but what's the minimum number of questions I can ask right are pegs participants answering over a thousand questions isn't ever going to be practical um for any amount of translation but but can we get down to sort of a minimal number of sets um of questions which is which is interesting I'll show you here just a couple of other results if you're used to seeing sort of an au or model fit in a different way um here's our example for type two diabetes again the purple is the polygenic score and all these others are our different poly exposure scores and you can see in lower GI pups you see even a stronger Gap um there between the polygenic scores and the poly social I think it's pressing showing it in one other way just to try to to drive um the results home here you can look at sort of prevalence plots so this actually is a number of cases it's it's um prevalence of that disease given the different percentile scores for the different scores so across a polygenic score polyal and poly exposure what I want you to see is that there's more discrimination you get a higher range from bottom to top there with the with the poly social and poly exposure scores then you do the polygenic scores so I'll stop there I know that was a whirlwind tour of things I've been thinking about with pegs so thank you um for your patience there I'm G to switch gears now to talk about the the partner research study um that Josh and all the leadership at all of us um has been just incredibly um supportive um and amazing on so talk about that a little bit um so stealing some slides from them I don't even how I stole it from but it's probably from you Josh um on sort of um priorities in all of us of the different um types of data that are they're integrated right there's a wealth of information in genomics and wearables the BIOS specimens you all have collected combined with a health record different behavior and especially fantastic social determinants of Health surveys and then a number of physical measurements how to sort of pull in more and more environmental um data I I find personally really exciting and doing it for all all the reasons um right sort of um pulling that sort of data in helps how do socioeconomic factors interact with exposures um for risk um to really think about about you know Health disparities think about risk and prevention what environmental biomarkers could identify risk for future disease could these sort of biomarkers be helpful in diagnostic what sort of exposure signatures do you find in patients with this or without this different condition um treatments and outcomes can we objectively obsess whether treatment or intervention might be effective right sort there's so many broad goals you could think about when you when you start thinking about pulling environmental data into all of us we had a fantastic Workshop um almost three two and a half almost three years ago right now where um all of us leadership folks from NS and people from the community at large um really spent a few days thinking about sort of priorities when you think about adding in environmental data and a number of themes um emerg sort of linking to geospatial estimates linking to climate and weather data exposomics whether you mean that from surveys or mass spe based sort of exposomics Technology thinking about personal exposure risk and and effective exposure on on diverse populations and we followed that up um with some really exciting um studies to incorporate um the environment into all of us um so really we sort of left that with sort of a a three-phase um approach to incorporating environment for phase one and two we really based on geospatial estimates and all of us has made just an absolutely incredible investment in getting addresses address histories for participants and then doing the geocoding um to link that data um to exposures they're doing that through a clad mechanism and we're hopefully having some collaborations with another niehs um effort that I'll just mention um briefly and then phase three is the partnered research study um that I'm leading here we're here we did a sort of a pilot case cohort study um looking at incident diabetes um so actually a ble to collect um sort of untargeted mass spec based exposomics and bios specimens from participants that had incident diabetes versus a a cohort um sample for comparison um which is really exciting we're in the middle of that data collection right now um I had a meeting yesterday we done 51 out of 71 um batches of collecting the exposomics data so it's moving along nicely oops I mentioned for the geospatial estimates um there's another effort at at ni um chords is the acronym the the acronym is under development again right now um but it's it's a um Court F funded um project to build tools to do geospatial linkages so it's taking data from NASA from EPA from satellites um in imagery and building tools to help people do those linkages um both to do software to build and curate web- based resources so let's say you're a human gen IST that's never done environmental Data before been there um and you say maybe I'm interested in air pollution it's building web tools and AI tools to say I'm interested in air pollution and it's going to pull up what resources are actually accessible to you here's the EPA and how they do it here's Noah and NASA and their versions right to to get you started um on adding environmental data into your studies and I said specifically um this all of us ancillary study um launched last July we're working on identify we identified participants um sent bios specimens to to a laboratory here in North Carolina that runs these Mass Spec um sort of assays and that that data is in progress now um so we'll have that on the case cohort study it will go into um the all of us data um Research Center and their ecosystem so that data will be available to everybody that has access at the right tier within all of us which is really exciting I said it here I'm sort of moving on phase three R said looking at sort of those that exposomics and and associations with outcomes you can imagine we'll be doing things like testing for associations of those exposomic sparkers with development of diabetes with the development of some of the diabetic complications retinopathy any of those others um with comorbidities and coort alties um and we'll have the wonderful opportunity to um sort of look across MMS um this is one of the largest exposomic studies ever done um so just having this number with this Geographic and demographic diversity is pretty unprecedented um and pretty exciting um so we're we're in the middle of that right now the masspec technology um that we've chosen it's the here lab it's it's here um up the street in Chapel Hill um collects a broad panel of metabolites um whether you want to call it metabolomics or exposomics so we'll have a number of um sort of endogenous um metabolites um representing host metab ISM so amino acids and polyamines and carnitine these sort of host metabolites then we'll have a lot that we can measure in the food um metabone um caffeine and um all sorts of um all sorts of different um sort of subcategories of vitamins minerals Etc a number of environmentally relevant metabolites posos phols parens Bates things you probably heard in in the news if nowhere else and then additional categories will'll be able to look at drugs and medicines folate vitamins um amino acids Etc this is what we're expecting um off there so we can be thinking about things that both use the annotated um metabolites that we'll be able to measure as well as thousands of sort of Quantified metabolites that don't have any annotation to them yet and this can support a wide variety of things so here's a very big slide with right now and all of us we're going to have genomic data exposomic data we'll have Society data societal data social determins of health and then phenotypes so you can imagine um hopefully some of the things I presented that we've done in pegs I I hope some of the similar thinking can come into all of us where we wherever we're thinking about xw whether that's from the biochemical data or the geospatial data can think about risk score analyses again using all of these different components um and to mention the the long read sequencing has epigenetic markers um for all of us too so can add that um as well can think about all the sorts of interaction analyses and then I think a lot about variance decompositions how much of this disease is due to genetics and environment and it's really exciting um like I said I've had a whole lot of fun in pegs and it is cute and adorable and being able to think about some of these approaches and something that has the mega sample size of all of us um is just really really exciting because that's that's what it's going to take um to to look at Gene environment interactions all of this and more um will be possible and I think that's really exciting and also I'm in it's biology I'm suddenly echoing um I don't know why but like said a lot of this is going to need to be fueled by methods development right there aren't tools out there to ask and answer all these questions um so there are a ton of really exciting data science um opportunities um as well that I I think are important so I'll stop there um again thank you for for dealing with this Whirlwind um I I I hope um it it was informative um and sort of the ways we're thinking about this of course I have to think all sorts of people and there are more people to thank um than just listed here but of course have folks with the the pegs leadership that I mentioned um I see somebody already out there from our external Advisory Board um even that I that I see in the audience I said people with the gis um all of our our colleagues at all of us um that are just amazing I said highlighting some of the the computational work some of the fromat assist that have done a lot of this um and our our collaborators and contract support and then first and foremost the the pegs and all of us participants that that give us generous access to the data so I'll pause there for questions and I'll stop sharing um there thank you I don't know if you could hear that but people were clapping for you you could probably at least see them clapping perhaps um so now it's open the questions I have as as people you can write them if you're online just write them into the Q&A and if you are um here you can just come up to the mic uh the first uh I I'll ask first one that I got online um and it was um how can we bring this knowledge to the clinics well I mean like we're we're in sort of the discovery phase right the the poly exposure score and things but I I hope that will actually directly be clinically applicable I I hope at some point um and there some huge and interesting efforts on adding social determinance and other exposures to electronic health records actually but I I'm hoping this sort of um information the sort of knowledge could help clinicians um learn better who to screen for what diseases or um you know which questions to ask and how to take that into more holistic um care so I know we're a long way from that um but but I hope we're less far away from that than than further if if that helps yeah it makes me think of the uh all the qu sections of our notes and and talking about social history how we could really expand that as well another one online um is uh do you have sufficient data for the vast numbers of variables to use deep learning methods to enable more agnostic results yes do we have the sample size right anytime you ask a statistician if they should have more samples we will say yes it's sort of like asking a dog if they're hungry like yeah the dog's hungry and I would love more samples um but we absolutely are able to use some of those those new machine Learners um and I've got some folks working on that I always think of um right anytime you're doing hypothesis generation I want to know what would be a useful hypothesis to follow up um so met like my dissertation 20 years ago involved neural Nets um I've been doing that sort of machine learning for a really long time um so you have to think there are cases where that sort of prediction that sort of multi-omic integration is is is really useful um and could could be sort of predictive power um so we've got some folks working in that space with that sort of machine learner and then I sometimes contrast that to the stuff that that maybe I was more heavy on here where we're trying to do hypothesis generation that we could we could follow up um right back to the clinical question we are long way away from somebody spending tens of thousands of dollars for multiomic in a clinical um setting so so we do we have one Avenue that I haven't talked about a ton with with machine learning and with methods to narrow down the search base for G by followup that's more targeted that that relies heavily on machine learning um but yeah we we can wield those tools um in this this sample size for sure great hi um thank you so much for your talk I have a maybe not completely formed question but um in looking at all of the awesome survey data that you guys have I think I first thought about how you get people to answer those surveys um maybe particularly people who live in areas that don't have great internet access or things like that because um I went to college in North Carolina and something we talked about a lot was the poor health outcomes in the western part of the state so just wanted to see if you could touch on that and maybe talk about any challenges you've had with like follow up with participants from those areas and things like that absolutely yeah it's absolutely a challenge so a lot of the recruitment was done with at at the beginning um when pegs was starting was done with very targeted recruitment drives um and so that involved things like Partnerships with church churches in those like rural areas that are a community that is trusted and that could interact um there were also a lot of joint recruitments in other initiatives um a lot of like women's health sort of initiatives so some of it we have more women um because women tend to volunteer for things but we also had some recruitments tied into some Women's Health um initiatives we had massive challenges during covid like a lot of people of people moving and relocating we do a lot of followup electronically now um right like we can take your addresses and like every adverti like there are plenty of you know computational resources that track who's where and and can reach back so we definitely had a gap in covid with people moving um and sort of reaching back out um as well for the most part um I'll make I'm gonna sound really flippant but I'm not these participants will just volunteer for anything like our participants are just amazing I don't know like I almost worry about them sometimes like you sure okay you'll keep going so also now we've switched a lot to electronic you know like we're entired electronic surveys um in sort of our ongoing followup um sort of callbacks um where we incentivized people to take the surveys whether it's with Amazon gift cards or or other sorts of of sort of simple things the other thing I will highlight about pegs that I haven't mentioned and just use this as an excuse is all of these participants are available to call back in our clinical Center um at niehs so pegs has also served as almost a core facility um for people that need uh bios specimens or or specific things we've had people do studies where we gotten we've called pegs participants back in to do skin biopsies on sun exposed versus unexposed regions and looked at mutation rates we've had participants call back for inhalation assays like they'll sit in a body pod and breee breathe diesel because we ask them um nicely so it's it's a really engaged community that cares a lot about it so in some ways I'm like I don't know the secret outside of like these people are just really really generous with their their time and energy awesome thank you another one online hi Allison great talk does pegs have disease incidence data are only prevalence are there efforts to bring in longitudinal data to pegs yes right now we have prevalence um it has been an ongoing effort but we're finally it's maturing we are linking to health records um first with UNCC Chapel Hill um so once we're there um we've we've done a lot of the the data sharing we've had a lot of IRB and reconsent um to get this done but I'm hoping in 202 will have electronic health records um and UNCC Chapel Hill doesn't just have their health system isn't just the hospital but like a large number of primary care um places are are included in that UNCC so it'll be it'll suddenly have longitudinal data sort of in the same way that all of us does um as we work to finalize that that health record so I'm very excited for that great great talk Allison bless beicker um question for you on how this really goes forward my you know head explodes with a number of variables you're talking about here and you know in of course in genetics it's taken us 15 to 20 years to go from GW was associations to people actually starting to work out cause and effect of these variants and it's been a really hard slog and I can't even begin to imagine what it that equivalent looks like for you from this poly exosome score to get to cause and effect or how what it is in that score that is disrupting the physiology that makes these people sick what what are you thinking about what that looks like yeah you know I understand and have no no easy solutions there right um and I I know you don't think we do um yeah like I I think some of the issues there there's some things that are way easier in genetics right like I know how to format that like I've got nice structured data it took us a while in the gws era for chips to evolve and get to you like there were challenges but I I can harmonize genetic data from the UK from Canada from all of us trivially by comparison of harmonizing environmental data um as well so there are all sorts of issues with sort of um data sharing in in exposomic and in Environmental Health Sciences um I I said we that pegs explore some of those things like for the most part if you want an epidemiological study there's so many data use committees and people have already like pinned what their little niche is um for this area so I think I think data sharing is going to have to be more expected for us to get there um a more open fisted approach that I think genetics has had more than Environmental Health Sciences and I don't triv trivialize that there are different issues to consider with that I'm not trivializing that that challenge um but I think that needs to happen um I think the opportunities for people in leadership at all of us in UK bio bank and other International efforts like a willingness to try to standardize at that level would go very very far um for that that like G by e or E is notoriously hard to find right for all the statistical challenges and all the high-dimensional challenges and honestly the mega sample size is what it's going to have to take to get there of just these huge mega samples within clever people doing causal inference in observational studies to help narrow that down and always really clever Partnerships between basic biologist mechanistic biologist toxicologist what yeah like it it's just going to be even more interdisciplinary to to get somewhere meaningful is is that helpful or sort of in line with what you're um it's very analogous to how we were what we were saying when jwa started like my God I I'm gonna take a prerogative to ask a follow-up question uh based on that that i' had written down which was it's interesting how similar the poly social and the you know poly exposure scores were in predicted value one being intervenable one not I was wondering if that's actually like just similar predictive values or it's actually the same similar for the same person and if so what does that actually mean does that mean that we tend to Cluster our exposures in some way you know that are the ones we intervene on not or you know is there something we can learn from that basically yeah we're just exploring this like some of this is hot off the presses um I would say it's more surprisingly the same predictive power than necessarily the same individuals um they are certainly correlated AB absolutely there is correlation but it is the predictive powers are more similar than I would have expected when I just looked at the correlations um but we we're really just starting to pick all of that um we're we're starting to disentangle this not just for diabetes but more more broadly yeah the predictive performance is stronger than I would expect on the correlation Neil I think you get the last question oh this one will be um relatively simple so if I heard you right you said that you were in 99 of the 100 counties what what is it about the onean county that you're not in I'm not going to remember the name of it I'm from North Carolina so I should it's actually the one to the furthest Northern right that is just really rural um it's like right at the edge of the Outer Banks there are not many people there and we want to recruit in that county for you so um but if I need to go to the beach if I need to go to the beach to to really make sure we're there I would I would make that sacrifice I would I would go to the beautiful beaches if I had to Great uh thanks for all the folks that ask questions online and it's also um it's been a great talk Allison I really appreciate uh you coming and giving this I think you give a a exposure to a kind of thing we don't get to talk a lot about um and the importance uh we often talk the importance of zip code and code but this is like a whole different layer on that as we talk about lifestyle environment and biology here so thank you very much for joining us thank you this is lovely and we're clapping one more time if you can't tell thank you thank you