Submind YouTube summaries
Thumbnail for 30min PhD thesis - Correlated Trait Locus (CTL) mapping

30min PhD thesis - Correlated Trait Locus (CTL) mapping

Watch on YouTube

Video summary

The speaker begins by reviewing the fundamentals of Quantitative Trait Locus (QTL) mapping, a method used to identify sections of DNA associated with variance in specific phenotypes like yield or tail length. In standard QTL analysis, researchers measure genotypes and single phenotypes across a population, using statistical tests such as regressions or t-tests at each genetic marker to find associations where the mean phenotype differs significantly between genotype groups. While effective for traits that show variation, this traditional approach has significant limitations: it cannot detect loci controlling binary traits like having hair versus no hair because every individual possesses the trait, and when two phenotypes are highly correlated—such as plant yield and susceptibility to infection—they produce nearly identical QTL profiles. This makes it impossible to distinguish between genes that simply increase size and those that specifically alter arm length or disease resistance, ultimately leading to a plateau in agricultural improvement where selecting for higher yields inadvertently increases vulnerability to pests. To overcome these constraints, the speaker introduces Correlated Trait Locus (CTL) mapping as an advanced method designed to analyze pairs of phenotypes simultaneously rather than looking at them individually. The core concept involves identifying genetic loci where the correlation between two highly correlated traits breaks down or changes significantly. Instead of calculating a standard effect size based on differences in means, CTL mapping calculates an "effect size" defined by the difference in correlation coefficients between different genotype groups (e.g., AA versus BB). By scanning the genome for these specific points where the relationship between phenotypes shifts, researchers can locate loci that decouple beneficial traits from detrimental ones. This allows breeders to select for high yield without simultaneously selecting for increased susceptibility, effectively breaking the genetic linkage that previously limited further progress in crop improvement and animal breeding. The methodology is demonstrated through both simulations and real-world data involving metabolic pathways in Arabidopsis thaliana. In a simulation using four phenotypes and six markers, CTL mapping generates matrices showing correlation differences at each locus, which are then converted into LOD scores to identify significant regions. A key advantage highlighted is the ability to reconstruct causal networks; for instance, when analyzing a linear pathway of three metabolites, standard QTL mapping might fail to detect associations for downstream compounds because no direct genetic variation exists in their means alone. However, CTL mapping reveals that while there may be no direct effect on the final compound's concentration, specific loci influence its correlation with upstream metabolites. This allows scientists to infer indirect effects and map out a complete causal network where one gene regulates an intermediate product which subsequently affects downstream traits, providing insights that single-phenotype analyses completely miss. In conclusion, CTL mapping represents a crucial evolution from traditional QTL analysis by shifting the focus from individual trait means to the relationships between multiple phenotypes. The speaker explains how this approach was instrumental in their PhD thesis and provides practical tools within R packages like `ctf` for handling complex crosses such as recombinant inbred lines or F2 populations. By utilizing permutation tests to establish significance thresholds, CTL mapping offers an unbiased, data-driven way to uncover genetic architectures that govern trait correlations. Ultimately, the speaker argues that while QTL mapping identifies direct effects on single traits, it is only a partial picture; understanding how genetic loci modify relationships between phenotypes is essential for overcoming biological plateaus and developing robust solutions in agriculture and biology where multiple correlated factors must be managed simultaneously.
Read the full video transcript
did want to talk a little bit about my phd thesis so back to some more like serious stuff that's not really serious like i just like talking about it so we talked a lot about qtl mapping right and that is just the association of a genetic marker with the mean in in different phenotypes right so you you have a marker in the genome where some individuals are one other individuals are two and then you look to see is there a difference in the mean of a certain phenotype right so when i was doing this presentation initially i just had like the definition right so a quantitative trait locus is a section of dna which is associated with variance in a phenotype a quantitative trait right so that's the definition of a qtl right we already saw this slide right so to detect the qtl in a population we need to have measured genotypes for example snips and a phenotype of interest such as tail length or yield or whatever you come up with right and then at each genetic marker you just do a regression or a t test to associate the phenotype with the genetic marker of interest so again here we go one by one oh this one goes automatically and doesn't have the little thing right so we go through the markers and at each marker we just say well okay is the a group bigger or smaller than the b group so that's kind of what qtl mapping is we do statistics to associate the marker and we show it as a lot score we already talked about that and then hand likelihoods are plotted for every marker on the chromosome we already saw this and then in rqtl i also showed you how to load the library so we load the library we load the data set and in in rqtl you have to scan one function to do a qtl scan if you don't want to do the t-test yourself or if you don't want to do the other test right and then you get you call the plot function and then the plot function generates this so there are some serious limitations in qtl mapping right one of the things that when i started my phd really was difficult for me to deal with is that ktl mapping only considers a single phenotype at a time right so we're only looking at yield or flowering time or some other phenotype which we associate and it requires the phenotype to show significant differences right if we want to qtl map the locus on the genome which controls the tail in ice we can do that for tail length right because every mouse has a slightly different length of tail but every mouse has a tail so we'll never be able to find a gene which controls if you have a tail or if you don't have a tail because every mouse has a tail the same thing holds for eyes right everyone has two eyes so doing a ktl mapping for for the number of well not so much the number of eyes because that might vary well not based on heritability of course but um so had the the phenotype needs to show significant differences and the problem is is that a lot of the phenotypes which are really really interesting do not show any difference right um every human has hair so figuring out where the locus is that controls hair or no hair is of course very hard to determine and one of the other big drawbacks in qtl mapping and we didn't talk about this before is that when you have two phenotypes which are very correlated to each other like the length of your arm versus the length of yourself then they result into very similar qtl profiles because the phenotype vectors are highly correlated to each other because the longer you are the longer your arms are the load side that it will come up with are very similar and of course that that that is genetically or like biologically that makes sense but for some things this is really annoying right because generally we we're interested in like where differences are controlled for and we also want to know for example where in the genome do we find the gene that controls if you have a tail or not so hit one phenotype at a time there needs to be significant differences and highly correlated phenotypes result in a very similar similar qtl profile which just means that you cannot distinguish load side that make you grow bigger from loci which make your arm grow longer so that's just a drawback so imagine two phenotypes which are very highly linked right so imagine that i am a plant biologist or i am a farmer and i am growing corn right and i'm very interested in the yield of my corn so the amount of grain that i get from a single plant and the susceptibility to infection then these two things are highly correlated to each other because bigger yields generally mean a higher susceptibility to infection and this is something that we have seen in many many phenotypes and here's one of these examples where we where people looked at wheat yields or the yields of wheat plants across the years right so you see here that well initially there was no real like um improvement right because we just did race selection and then we did pedigree selection to kind of improve our plans but when we started like in 1955 using scientific breeding and especially in the 1980s with the advent of qtl mapping people figured out which locations on the genome of this plant controlled the yield and we've been doing this now for like 20 years and the yields are decreasing right you see that in the beginning when we started qtl mapping we had a very high or very rapid improvement in the amount of food that we could make out of a square acre of a field but this has been more or less stabilizing since like 2000 2005. so we found the initial outside that were controlling the phenotype we selected the animals for having positive load psi but in the end we we selected this a couple of times and now we're in a situation where our plants are more or less optimal and there's this kind of improving plant yield further also makes them more susceptible which makes that the yield goes down again and this is this is a very serious problem in science and in production as well right because the way that we as humans live on this planet is that every year we need more resources because there's more humans so we want this increase that we get from the scientific breeding approach or more or less the selection approach using qtl mapping and these kinds of things to continue improving but we can't because the more we improve our plans the more susceptible they become the less yield we get in the end right so at the beginning of my phd i thought long and hard about this and i said that no we should find a different method right because we are interested in these two things at the same time the susceptibility on the one hand and the yield on the other hand and these two things are coupled together so we want to instead of map the differences in mean we would want to map where this correlation is breaking right because if we find a genetic locus where the correlation breaks then we can actually select for this locus and then we can continue improving without suffering the the effect from the susceptibility so that's why i define ctl mapping as a correlated trait locus ctl which is very poorly chosen in retrospect because ctl also stands for cytotoxic t cell and there's a lot of literature about cytotoxic t cells and their literature amount is improving so no one can find my method so everyone who search for ctl mapping on google gets like genome-wide associations where people use cytotoxic toxic t-cells so it's very poorly chosen but a correlated trait locus is a section of the dna so a locus in the dna which is associated with differences in correlation between phenotypes so ctl mapping is very similar to ktl mapping it's just the difference is that it's multi-phenotype so instead of looking at a single phenotype at a time we're looking at pairs of phenotypes and what we want to do is we want to identify genetic regions where there's a difference in the phenotype to phenotype correlation conditional on the genetic marker that we're currently looking at and of course it's an unbiased data-driven method with no prior information so there's no type of bias that flows into it right the same thing as qtl mapping we just measure the genotypes we measure the phenotypes and then we do the association analysis and ctl mapping is similar in that sense so there's no like iterative process or these kinds of things involved so the idea that i when i came up with the idea is what that ctl mapping should be applied in classical selection in breeding to improve the economically interesting phenotypes so the the idea was that you select for a beneficial ctl locus similar to qtl so in the case of yield and susceptibility you want to break them right so you don't you want a locus at which these two phenotypes are not showing any correlation because if they are showing correlation and you're selecting for it then the next generation will still show correlation right and the idea is is that by selecting for this beneficial load psi where the correlation breaks you can you can break the linkage so the correlation between these phenotypes in the next generation when we started doing this we also figured out that when we combine correlated trait load side together with quantitative trait load side we can build kind of a phenotype by phenotype causal network and i'm not going to talk much about that but i want to talk about this initial idea that by trying to find loci at which two phenotypes which are normally highly correlated are now not showing any correlation and then selecting for that locus will be able to break this this linkage um so that's the thing right so the yield plateau tells us that both phenotypes yield and susceptibility because they are highly correlated with each other you will get the exact same qtl profile right so when you select for high yield you will also select for increasing susceptibility which you don't want because that means that you have to use more and more chemicals to get rid of all of the bugs every time so the idea is that ctl mapping finds load cyber correlation between yield and susceptibility is lost we can then use the ctl information to break the correlation between yield and susceptibility so how does this look so head this was one of my figures from my thesis where i just said well head this is a simulation and so if we have phenotype a which is the yield of the plant we have phenotype b which is the susceptibility we have the aa genotypes which show strong correlation right the same correlation that we see overall in the population and we have here the bb genotype which shows low to no correlation which you can see that there's no like it's just a cloud of points and there's no straight line right so if you see a picture like this which genotype should we breed should we breed the aaa individuals or should we breed the bb individuals so that's a question to you guys in chat if you're still listening right so that was the whole idea behind it right that if you have two phenotypes which are highly correlated overall at each marker you look at the correlation between the individuals which are a a and the individuals which are bb and then you will find if you're lucky a locus in the genome where these two highly correlated phenotypes are not correlated so in this case of course if you want to breed the next generation you want to breed individuals which are bb at this locus right because the bb shows no correlation and in this case we have a positive phenotype linked with a negative phenotype and we can also have like a negative phenotype length with a negative one and then of course you would probably want to have the correlation and show the lower tail right so select the ctl genotype to unlink two phenotypes and in the next generation which should show a decrease of correlation between the phenotypes and this allows again in the next generation to select high yield without increasing the susceptibility right and then in in the next generation you would just select individuals based on qtl information like you would normally do so the methodology here i explain it using recommend name red lines there is a package on chrome which actually is called ctl and this package handles many more complex crosses so i i just explain it here using aa and bb individuals but it also works if you have an f2 population where you have a a a b and b b individuals so recombinant in red lines um i'm i'm at the example here is assuming that i have four phenotypes which have been measured and i have six genetic markers right so i have four different phenotypes and six genetic markers and um at every locus right at every marker there an individual can have either a a or bb just to simplify it for the for the presentation so the way that this methodology works is first you select a phenotype called p1 right so that's your first phenotype then you select a genetic marker so the first one right you split the individuals into two groups by their genotype so you have a group of individuals at the fir which at the first marker is aa and then you have a group of individuals which at the first marker is bb so here you have a population for example of 100 animals and 50 of them go into group 1 50 of them are in group 2 because these 50 faa these 50 fbb then what you do is you calculate your correlation for both the aa and the bb genotype what you do is you do p1 times all of the phenotype correlation vector right so you just say well calculate the correlation of your phenotype p1 with the phenotypes which are showing a a right and then you just get a vector so phenotype one of course shows a correlation of one to phenotype one because two things which are equal always have a correlation of one p1 and p2 at this marker show 0.1 uh correlation um p1 and p3 0.5 p4 0.8 right so here p1 and p4 are highly correlated at this marker you do the same thing for the bb individuals right so you again get a vector with four correlation coefficients um of course p1 versus p1 is still one um and for the other three phenotypes you also get a correlation so correlation in pop in the aaa individuals correlation in the bb individuals so then the next step is just to define the effect size right so the effect size in in qtl is defined as the difference in the mean between the aa group and the bb group but in in ctl mapping the ctl defect size is calculated or defined as the difference in correlation between the aa and the bb groups right just like qtl but now instead of looking at the difference in mean we're looking at the difference in correlation between phenotype right and to make it easy we're just going to take the absolute difference right so not negative and positive we're just going to say that um like the the absolute difference right so minus 0.1 becomes 0.1 so when we do that so here we have the two vectors from before and now we just calculate the difference vector so of course p1 because the correlation in of b1 to p1 is of course one and one the difference is zero the difference from p1 in aa to the p1 to p2 and vb um is 0.1 0.2 right so we just take the difference so we just subtract these two vectors from each other then the next step of course is now we have mapped one marker is to do all of the markers right so what i do is i take the vector that i just had and just put it on its side right so now here we have the markers in the in the columns and we have the different phenotypes in the rows right so this is the result from the last slide you can check that it's 0.00.10.20.7 and indeed here you see the same thing right and then this is the difference vector for matrix 2 different vector for marker 3 difference vector for marker 4 and so on and like i told you we assume that we only have 6. right so i repeat this calculation for every genetic marker so multiple difference vector for our selected phenotype of course that when we map p1 against p1 it will all yield a difference of zero right so we could have not mapped this and just skipped it but that just for completeness sake i just want to show you the whole whole matrix let me get a sip of one sorry it's been a long lecture all right so here we map p1 against p1 always a zero and of course for the other phenotypes we don't get that so now we need to of course find what is significant right is this difference of 0.6 in correlation is that a significant difference so what we do is we repeat the same thing at 10 000 times just like i showed you guys for qtl is we break the link between genotype and phenotype right so in this case we're just assigning genotype vectors at random to the individuals just like we did before in qtl mapping but in qtl mapping we assigned the phenotypes randomly right but now since we have two phenotypes we want an individual so the two phenotypes of the individual to stay the same but now we just give it a random genotype vector in the end right so we redo the whole analysis we remember the maximum score and then we make a distribution out it and then we find our five percent and one percent thresholds for significance values right it's just the same way so we're just going to permute ourselves out of problems by just saying i'm just gonna instead of assigning a new phenotypic value for each individual based on the back that we have we're now just going to assign like a new genotype so ctl uses 10 000 plus permutations to add sign significant of course in the package because i did study this stuff for four years i also devised a method to directly calculate your p-value using mathematics so of course here we then convert so we convert the differences in correlation that we see to probability values so how likely is it that there is a real correlation difference at that point and then we do the next step which is just saying the same as qtl where we convert the p-value to the lot score so we just take the minus log 10 of the p-value so by converting this have we performed qtl mapping for p1 as well so because we we have the data anyway so we have p1 which we can ctl map against the four phenotypes that we have but we can of course also just do the qtl mapping right so for p1 we get four vectors of lot scores from ctl mapping right so every every every every correlation difference is transformed into a lot score um and beside that we have the information about p1 itself right so because we can also associate p1 with every one of the six markers that we have so how does this then look so this is the way that we visualize it so we take our qtl curve of p1 and we just plot it on the top and then we take the ctf curves of p1 versus the other phenotypes which we see on the bottom right so on the bottom is the lot score and the negative lot score of the ctl score and of course the on the bottom we then see 4 or five or six or seven lines no matter or how many phenotypes we had all right so that was the ctl mapping method it's a relatively easy thing right and in the end we find loci in the genome where we see that there's for example a qtl controlling the variation in p1 but we also see that at this locus p1 loses correlation with some of the other phenotypes let me actually pull up one of the other presentations that i did about ctl mapping where we use some real data right just to show you guys how we can use this information it should be somewhere in presentations ctc transmission ratio distortion here so let me open this up and all right can i actually easily swap that just swap the powerpoint yes i can so let's go to properties and then switch to this one right so what we were looking at here is in arbitropsystaliana we have metabolites and these metabolites are known to be in a linear pathway so we have something called hydroxypropule which is then using an enzyme is transformed into material sulfoneupropyl and then using another enzyme this is transformed into material thiopropyl right so it's just three metabolites with two enzymes in the middle there is a major regulator of this pathway on chromosome five and all of this was known right so then i did the ctl mapping right so the first plot that i'm going to show you is and remember the colors right so it's it's green red orange right so green is on the top of the network then green gets transformed into red and red gets transformed into orange so here we see this the qtl profile of hydroxypropyl on the top right so we see that there's a major regulator of the difference in hydroxypropyl and chromosome 5. the same thing holds for chromosome 4 there's a marker on chromosome 4 which also controls the heteroxypropyl concentration in in the plan but then we start seeing that if we do the ctl mapping with material sophomore and material theoprophyl right so the red one and the orange one we see that we do find a little locus on chromosome one and that is strange right because we never had any indication from qtl mapping that something on chromosome one was actually driving the hydroxypropel region but we do get the idea that well there is something on one which makes which which makes hydroxypropyl lose its regulation or lose its correlation with the other two phenotypes that we're looking at and we see the same thing on chromosome 5 right so on chromosome 5 we learn nothing new because we already knew that there was a major driver of this network when we then look at the middle phenotype in the pathway of the middle metabolite in the pathway methyl sulfaneopropyl what we see is now hey that's interesting there is a little qtl on chromosome 1 for the middle phenotype so there is something on chromosome 1 which is controlling the concentration of the the middle phenotype of course there's also something on chromosome 5 which is kind of passed down right from the initial one so the in the initial concentration of hydroxypropule is transformed into materials sulfoneupropyl right so the more you have at the beginning of course the more you have from the intermediate product as well so here we also see the um the ctl line and we did indeed see that that at this locus we do get the idea that yes no there is correlation between the amounts of sophomore of uh hydroxypropyl sulfanium propel and the theoprovium so when we then look at theopropyl now we see something interesting because if we look at the top we see that when we do the qtl mapping of this phenotype we don't know where the concentration of this this this um metabolite is controlled from there are no significant regions so if i would do an experiment using material theo material theopropel measurements right and it would scan across the genome i would learn that there is no locus that is controlling the concentration of material theopropyl however if we look at the ctl map we do see that we get significant load side so we do see that the the the method tells us there is something on chromosome 5 which is influencing the correlation of theopropyl hydroxypropyl and materials materials sulfoneopropyl right and the same thing again on chromosome 1. so what we see is that for the first two phenotypes we don't really learn anything new except for that there is something going on on chromosome 1 for hydroxypropel we learn for the middle one we don't learn anything or not that much because we already knew that there was something on chromosome 5 controlling it but for the last one we don't get any qtl and because we get no qtl as a geneticist you're stuck here you cannot say what's happening to teopu but we can from qt or from ctl mapping we learn that no there you if you want to influence this phenotype you have to be on chromosome 5 and there might be something on chromosome 1 which can also influence it so the idea is is that we have this network right so this network is driven so if we combine the qtl information that we get so on chromosome 5 we see that the strongest association with chromosome 5 is with hydroxypropyl then the the lot score then because this the concentration of this one influences the concentration of this one so we can still see the effect of the chromosome 5 locus on this one but we learned that this effect of chromosome 5 is not a direct effect it is an indirect effect it goes via hydroxy so the thing on chromosome 5 is not directly influencing material sulfur neutropia it is influencing hydroxypropel and that in turn is influencing material sulfone protein so the exact same thing have we then find for this locus here we don't find any direct association of chromosome 5 with this phenotype but based on the ctl and the strength of the ctl right because the thickness of the line determines the strength we now learn and we now start seeing that indeed this is kind of a linear network right because we see that the red metabolite should be in between the two because there is a very strong ctl from the red one from the green one to the red one but from the green one to the yellow one it is much lower and from the red one to the yellow one it's also lower but it's still detectable right so we can start building up this causal network and seeing that indeed the network should be hydroxypropyl causes changes in sulfoneupropyl which then causes changes in material to profile and without ctl mapping we would have never looked at chromosome 5 for this phenotype because there is no direct association only based on the correlation can you see that the correlation is lost at that point good so that's the whole big idea that's what people gave me my doctor title for seems not much but was a lot of work was a lot of work to uh to to do all of this and uh thank you guys for uh actually staying until the end right so you can see that i actually use the same phenotype in this presentation as well but the idea is is that qtl mapping is only a part of the puzzle qtl mapping gives you direct effects on the means of phenotype but in the end it's not about single phenotypes it's about the relationship between phenotypes and how genetic loci kind of modify these relationships good so that was what i wanted to tell you today so for today just as a quick overview we did phenotypes heritability i tried to explain to you guys what qtl mapping is and how you need to use experimental crosses i told you guys about gwas i didn't tell you about the bfmi it will be in the slides that i upload but i just skipped that part and i will talk to you i talk to you about ctl mapping also the fine mapping part is not directly in the in the presentation because i skipped it today because we were kind of running out of time because we took an hour for the assignments all right so for me that's it for today if there's any questions remarks other things then then please let me know throw it in the chat i think all of the four guys that made it to the end like you're amazing thank you for being here and of course for my moderator who's also probably still here um bacon misha thank you guys for joining and staying until the end um that was it for me it's uh five so people on youtube um see you on the flip side so see you next time