Submind YouTube summaries
Thumbnail for The HIDDEN viruses in our DNA could unlock HIV Cure!

The HIDDEN viruses in our DNA could unlock HIV Cure!

Watch on YouTube

Video summary

The human genome contains approximately 8% endogenous retroviruses (ERVs), which are essentially fossilized remnants of ancient viral infections that integrated into our germline millions of years ago and were passed down through evolution rather than resulting from recent infections. Unlike active viruses such as SARS-CoV-2, HIV is a retrovirus capable of reverse transcription, converting its RNA into DNA to integrate permanently into host chromosomes; however, this integration only becomes heritable if it occurs in sperm or egg cells. While much of this non-coding material was once dismissed as "junk DNA," research has revealed that these elements are not merely inert debris but have been repurposed by nature to play critical functional roles, a concept supported by the discovery of "jumping genes" and the study of ancient viral fossils in fields like paleovirology. A prime example of this evolutionary co-option is the *syncytin* protein, which is essential for placental development in all placental mammals. This vital protein was originally derived from an ancient retroviral envelope gene responsible for cell fusion, demonstrating how viruses captured by our ancestors were transformed to enable extended pregnancies and complex reproductive systems. Beyond reproduction, these viral sequences have also been harnessed to boost human immunity; regulatory sequences like Long Terminal Repeats now help control antiviral genes, while viral envelope proteins have evolved into agents that block infections from related viruses. Although reactivation of these dormant elements can occasionally contribute to diseases like cancer or neurodegeneration, their overall impact highlights a history where viruses drove significant genetic innovation rather than just causing harm. Building on the understanding that nature has already solved complex problems using viral mechanisms, scientists are now exploring how to apply these ancient strategies to cure HIV. An NIH-funded project aims to implement a "block, lock, and stop" strategy by focusing on silencing latent HIV reservoirs within T cells, which currently prevent a complete cure. Researchers are investigating KRAB zinc finger proteins, cellular tools discovered recently that physically bind to and silence endogenous retroviruses, with the goal of engineering new repressors to permanently lock HIV in a dormant state. By mapping how these natural silencing mechanisms function in human immune cells, scientists hope to borrow from nature's own viral defense systems to finally eradicate the hidden virus hiding within our DNA.
Read the full video transcript
Hey everyone, Raif Derrazi here, and today I'm excited to have a conversation with our very special guest, Dr. Cedric Fechot, who is one of the scientists or investigators in the HOPE collaboratory, which I've talked about numerous times on this channel. He has his own lab with investigators working under him. Today, he's going to explain what HERVs and ERVs are, and the role they play in our body, the impact that HIV may have on them, and then at the end, briefly cover how they may inform an approach to HIV cure research. I will disclaimer a little bit. This is going to be a more in-depth conversation presentation about HIV research. It gets a little dense, a little heavy at times, but we're going to try to go through it slowly and try to explain things in great detail. I'm going to ask questions that hopefully will help make it clearer for you or help to summarize them in terms that maybe you'll understand better. If you don't get everything, if you're still missing things, don't get upset or dismayed. If you don't get everything right away, if something's, you know, fly over your head, that's okay. The point is that you're starting to pick up on some more things related to science and research and learning a little bit as we go along, and then the more I have these kinds of presentations and these kinds of talks, hopefully you'll begin to pick up more and more pieces, and then the puzzle will start to come together and the light bulb will go off and you'll start to understand things. Because in the end, my goal with all of this is that you have a sense of ownership in your journey with HIV and also feeling connected and part of the research that's going on towards an HIV cure, towards an HIV vaccine, and that you're able to follow along somewhat when news comes out or new studies are released, things like that, so that you're not an outsider just in the community, but you are a part of this process and this investigation and all this research into HIV. Hi everyone, Raif Derrazi here, and today I'm excited to have a conversation with our very special guest, Dr. Cedric Feschotte. But first, his biography. Cedric Feschotte, PhD, is the Barbara McClintock Professor of Molecular Biology and Genetics at Cornell University. His laboratory studies the evolution and biological impact of mobile genetic elements and endogenous viruses in a wide range of eukaryotes, including humans. And folks, if you don't get what all these things are, that's okay. We'll we'll cover some of them at least. Dr. Feschotte obtained his bachelor's degree from the University of Toulouse, France, in 1996 and completed his doctoral studies in 2001 at the University of Paris, working on mosquito transposable elements with Professor Claude Mouchès. From 2002 to 2004, he was a postdoctoral fellow with Dr. Susan Wessler at the University of Georgia in Athens, Georgia, where he investigated the origin and amplification mechanism of plant transposons. He launched his independent laboratory in 2004 as an assistant professor at the University of Texas, Arlington. He then joined the University of Utah School of Medicine in 2012 as an associate professor in the Department of Human Genetics, being promoted to professor in 2016. In 2017, Dr. Feschotte relocated his laboratory to the Department of Molecular Biology and Genetics at Cornell University. He received the Empire Innovation Award from the state of New York in 2017 and was elected fellow of the American Association for the Advancement of Science in 2019. In 2023, he was appointed as the Barbara McClintock Professor of Molecular Biology and Genetics at Cornell University. Cedric, it's so fantastic to have you on the channel. How are you doing this morning? >> Doing great. It's great to be with you, Raif. >> And I'm curious. Um it says you're appointed as the Barbara McClintock Professor. What does that mean? >> [laughter] >> Well, you know, that's one of those uh endowed professorships, but I have to say, this one means a lot to me because it's named after uh the person, Barbara McClintock who discovered transposable elements, the thing that I study. And it's actually kind of coincidence. Uh, it's not it was the title was not created for me. It just was in existence at Cornell because Barbara McClintock is an alumni of Cornell University. She was both an undergraduate and a graduate student here. So they had that title available and it was, you know, given to me. So it's really extra special to be called the Barbara McClintock Professor uh, for me. Yeah. >> Yeah, that's super special. And um, you clearly have a very uh, storied career, so much experience. So I'm really excited to have you on. Um, folks, today's talk is a little different in that instead of interviewing Dr. Feschotte like I would normally do, he's going to be presenting and explaining what ERVs are and how they ultimately relate to a potential HIV cure. And I get the opportunity to interject and ask questions as we go along. Hopefully I'll I'll ask the questions that you at home would be asking. Many of you have expressed interest in learning and having these kinds of talks. This way you can have a better understanding of the types of HIV cure research that are happening here and around the world. And without further ado, Cedric, I'll let you take the lead from here. >> Cool. Thank you, Raif. Yeah, super excited to do this with you today. Um, all right, let's see how that goes. Um, so this So yeah, I'm just it's we're going to be really like just covering like really background and introduction on what these things are, these endogenous retroviruses. I like to think about those as the viruses that we all have uh, within us, in all of us. Uh, all humans have these retroviruses in our breeding to our own DNA, right? Um, so just uh, I thought I would start with just a definition of what we mean by the genome. Um, so I did sell this slide this lost its animation, but it's it's it looks complicated, but it is complicated, you know, that's the recipe for life. That's the blueprint to make an organism. Every one of your cells in your body as DNA in it. And this DNA, together, we call this the genome, it's contained within the nucleus of a cell. And they it's partitioned into molecules called chromosomes. Each each of which is made of DNA and they basically encode again the sort of blueprint for making an organism really in the form of of genes is what we are all familiar with. These are the genes that of course inherited to the next into the next generation. These genes, you know, just a few terms because we might be using these terms. They the genes actually uh code for an intermediate molecule called RNA, which is related to DNA, but it's it's a slightly different chemically speaking. So, they are what we call transcribed into RNA. In every cell, there's this stuff happening all the time as we speak. And then these RNAs typically get translated, we we say, into proteins. And together, the proteins and RNA really make very assemble in very sophisticated and complicated machines that basically drive the function, the development, the physiology of our cells and and and of tissues and organs and and and the making of an entire organism. Yeah. So, that's what we call the the genome. >> So, is is saying genome and DNA, is that kind of interchangeable or is there a difference? Cuz I always kind of was confused when I heard genome and then DNA and if those are different or the same? >> Yeah, very much. You can think about genome and DNA. Now, there is organisms or creatures out there, things out there that have genomes that are made not of DNA but RNA, it turns out. In fact, in fact, viruses, many viruses like SARS-CoV-2 or HIV have an what do we call it RNA genome. So, that's their their primary molecule for life. But, as you know, HIV has an additional kind of trick where it makes DNA from its RNA. And which is kind of a reverse of what I just explained here, which is why we call these retroviruses because they want to reverse their RNA back into DNA. And as we're going to talk about in a minute, this DNA can then integrate in the genome of the host cell, which is then DNA, right? And then make more RNA of it and then more and then more viruses, right? So, this is actually retroviruses like HIV have a really complex, you know, lifestyle, replication cycle. And when they go from RNA to DNA and then RNA again and then they get packaged. So, you know, see what's containing with the within a baby virus is RNA, not DNA. So, in a way they have an RNA genome. That's what virologists talk about. SARS-CoV-2, which, you know, we all know about now, the COVID, you know, virus has an RNA genome and this one never makes DNA. So, actually it makes a copy of the RNA directly from the RNA. So, it has only an RNA genome and we call these RNA viruses. Yeah, but generally speaking, you know, in organisms like multicellular organisms or bacteria as well or single, you know, single-cell organisms like yeast or whatever, all the genomes are made of DNA. Yeah. >> And so, because HIV has this ability to for reverse transcription? >> Yeah. >> Um then it's able to become part of our DNA. That's what makes it so >> Exactly. >> difficult as opposed to COVID. COVID doesn't do that, right? It doesn't become part of our DNA. >> Yeah. And I I'm again again I I'll get into these in a just in a couple of slides from here. Yep, exactly. Now, today we're going to be speaking about the human genome. Um, but what I'm going to tell you about the human genome, I just want to say is actually very uh it's really common to basically all all animals I would say to some extent. Um, it's nothing really that special about the human genome, but today we focus on the human genome. So, um, you know, so about what? 40 more than 40 years ago now. Um, now 30 years ago something whatever. In 1990 it was a project called the human genome project that was launched. The goal of this project was to sequence, meaning to decode all the letters of the DNA that make up the human genome. Uh, that's a that's you know, at the time it was a really daunting task uh, because the sequencing technology, the way that we have to read these DNA uh, was very very slow and very costly. Um, now there is let me tell you like so for I just go back here for a sec, you know, this is the these are the letters that make up DNA. As as uh, you know, most of the viewers would probably know there's only this is an alphabet, but there are only four letters. So, in in a sense it shouldn't be that complicated, but there's still many combinations that you can make of these four letters depending on the size of the genome. The human genome is made up of three billion base pair or nucleotide sequence. So, this A, T, G, and C. Uh, 3 billion. Okay, so that's a very long uh, sentence, you know? So, it's a very big So, it's kind of like instead of a sentence, by the way, you can think about it as a whole book, a big encyclopedia, right? And uh it's And it's broken up into like words and sentences. We can think about this as the genes, you know, in a sense. They have sense. They They They mean something. And so, the goal here with the Human Genome Project was to really simply read, to begin with, to know what the book tells tells us. We need to read it, right? So, that was just simply to read uh literally this sequence for the 3 billion base pair. We call these base pairs because they make base pairs into DNA. Um these 3 million these 3 billion letters we have to read them. And that actually was very hard because at the time you could only read a few hundred at a time into like one reaction chemical reaction in the lab. So, it took like forever, and it was extremely costly. It wasn't It took 20 years. Uh I really I'm getting all of my You can edit this. I'm getting all of my dates right uh wrong. >> [laughter] >> It took more than 10 years. I'm sorry. Yeah. It took more than 10 years to um to achieve that, although it was still partial. And the first draft of the human genome sequence was released in 2001 in a couple of papers. And it cost $3 billion to get there. And I think that's probably an underestimate. So, this I bring this up because today 20 years later we actually can sequence a human genome in a single lab. I can do this in my lab. And it would cost us less than a thousand bucks. Okay. We wouldn't get really maybe high precision. There would be some mistakes here and there. It would probably not be complete, but, you know, just give you an idea of the how the cost of sequencing DNA has dropped in the last 20 years. I mean, it's just like phenomenal. Uh the advances the technological advances in the way we do this now, sequencing DNA is completely different than the the way this first draft uh was produced back back in 2001 and back in the '90s and 2000 early 2000. Yeah. Um I also wanted to say something else. Um Um Yeah, so uh no, I don't remember. Oh, yeah, so I wanted to What I wanted to say is like I talk about the human genome sequence, but you may wonder, you know, which humans are we talking about? Which Who was being sequenced, right? Which >> We sell. >> Okay. Uh well, um that's a hard kind of a hard question, but to simplify the answer, it was an anonymous genome, I would say. It was actually a composite genome that came from different individuals. And of course, the identity of these people was never revealed. Uh but it was meant to be a generic, what we call a reference human genome. Okay, it wasn't It's an anonymous genome. Uh as we're going to talk about a little bit, I think maybe later. Of course, human genomes are different. You know, if you compare the genetic makeup of two different individuals, of course, that's why we're all different as humans, you know. Um we are all very different, but remarkably, the human genome is actually There's not a lot of variation from one individual to another. There's very remarkably low variation compared to other organisms like say fruit flies are much more variable between individuals. Humans have gone through like a very serious bottlenecks during human history. Um and as a result, and we all came out of Africa, you know, not too long ago. And as a result, the diversity the genetic diversity in humans is actually very very low compared to say you know, like I say flies or even like chimpanzees, um which are our closest relative to humans as far as animal species goes. Uh chimpanzees are actually more diverse human gene uh genomes chimpanzee genomes than than humans. But anyway, so so when I talk about the human genome, you have to think about it as a generic, average human genome. But there are there is definitely variation, of course. And that's another, of course, uh focus of studies is to catalog. Nowadays, because it's much cheaper and faster to sequence human genomes, we have access to, you know, hundreds, thousands of human genome sequences that we can compare one another. And I think it's fair, probably, uh it's safe to say that, you know, down the line, in the near future, probably everyone would be born with their genome sequence completely completely read, completely sequenced. Yeah. But we're not there yet. >> would it be to to have your health care journey start with sequencing your genome? >> Yeah, exactly. Well, you know, there is, of course, a number of potential ethical issues about this, obviously. Uh um as you know, you know, you can you you can pay and get your genome sequenced. Like, there are many companies, like most famous probably called like 23andMe, that would do that for you and would tell you then would send you a report. Now, we I don't know if they send you back the sequence. I don't think they do. But they would send you a a report of, like, you know, your likelihood of having certain susceptibility to disease, which, of course, can be very, very important in the prevention of those diseases. Uh we all know that you can do um you know, there's a number of genetic testing that's being done already in very routine basis, right? But yeah, um I think it's it's I think it's very likely and and uh that that genome sequencing would be done, you know, like kind of systematically, if you wanted to, of course, if you know, you could opt out, probably, for this. Um yeah. Okay. Should I move on? >> Yes. >> Okay. Yes, and so, actually, just only about 2 years ago. It's kind of remarkable that it took that long, right? It took like It took 10 years to get and I told you it was very laborious and of course it involved like it was a multi It was an international effort to do this. But it took, you know, 10 years to get the first draft, but it took another 20 years to get to something that's actually we know is has no gaps. It's complete from what we call end-to-end chromosomes. All the chromosomes are complete and there's no gap. And I I bring this up because um uh the reason why that affected our work is like the piece, the parts, the gaps that were missing in that draft, you know, that took so long were actually these complex repetitive sequences that we study in our genome including some of these endogenous retroviruses. It turns out these are the most difficult part of the genome to to sequence uh because they're repetitive and it's a little bit like a puzzle, you know, if you try to do the sky of the puzzle the blue, you know, it's all the it's all the same color. It's kind of usually what you keep for the end. I don't know about you, but I keep this for the end. This is like the hardest part, right? That's exactly like this. >> That's a good analogy. >> Yeah, and it very much is the same problem that computers face to put to to put it together, to put the puzzle together. But it can be done and it's been done now very very accurately for for again the reference human genome. Uh again, to get to that completion, that level of completion, it takes it takes a lot of work actually. All right, so one of the really striking things that was learned when the first draft was released. Of course, you know, you have to read the genome and then you have to try to make sense out of it. And you know, we don't understand the language. That's the thing. The goal of sequencing the human genome is to try to understand the the lexicon, you know, the language that by which, you know, the cells functions essentially in life works, right? So it's a big undertaking. Um but I think one of the most striking observation, which was already known honestly prior to the sequencing of the human genome, but then it became very clear when you got the sequencing in front of your face, start to analyze the sequence, is that the amount of DNA, the amount of the genome that ends up coding for the things that you would think matter the most, which are the proteins, you know, that that that are coming from the the DNA gets into tran- transcribed into RNA and then RNA gets translated into proteins. This is These are the workhorses of the cell, right? They do pretty much everything. But this uh DNA only occupies 1% about 1% of the human genome. So I'm telling you like 99% of our genome does not eventually encode for the cellular proteins that make our cells and makes of makes us live and reproduce and react and and and and protect us against viruses and so on, right? So it's really a tiny fraction. >> This was really kind of shocking to me when I saw this and really interesting. Um I'm curious though that 1% is that what dis- what determines like all of our physical attributes, everything physical about us? >> Yes, a gre- great question. Uh so that's a huge question. How much of the 99% is actually important, right? At the end to to for the cells to function. Uh how much of the human genome is functional? Certainly >> Yeah. >> uh it's not just the 1%. So that we know for sure. >> Okay. >> Because a lot of the other DNA here is what we call non-coding DNA, but it includes sequences that are responsible for turning on and off the genes. And we know that's extremely important for development, for how do you react to any stimuli. You have to turn on and off different genes, and that's why our cells like neurons will be very different than liver cells because they are different sets of genes that are being turned on and off. Not every gene is on at every given time. Only a a small subset of genes need to be always expressed make proteins. Those are like we call housekeeping proteins. It's a minority actually of the proteins. So, the rest is what makes different cells different. Uh different cell types different and different organs do different things. And these are the majority of the protein coding genes. But this is the regu- the regulation then of the genes of course is absolutely critical to understand the development and the life of an organism. And these sequences that turn on and off genes are part of the non-coding genome. So, you know, sometimes people talk about this non-coding genome as sort of the dark matter of the genome. Again, because we still don't know how much of that part of the genome is important. So, what is what are the estimates out there? Well, I think most uh genome biologists would agree that at least about 5% of this non-coding genome is critical, at least. But that still leaves a lot of DNA that we still don't know the function of, right? The rest is much much less clear. It could be a lot more, but no one has done the experiments to remove chunks of these and see what happened in cells actually. It has not been done yet, even though we have technologies now like CRISPR, you may have covered before. Um that that enables us to do these very precise removal experiments. But people haven't done this at that kind of scale. So, we still don't know. It's It's a really important question in biology is like how much of the human genome truly matters at the end. And maybe the viewers would wonder, why would you have DNA that don't doesn't even matter? All right? It doesn't It wouldn't seem to make no sense. >> were going to say the percentage the estimate, I was thinking something over 50%. It's Oh, I'm thinking that's reasonable to assume that something over 50% is you know, usable. Important, but it's 5, which is a lot smaller. >> Yeah, we have only the in the like 10% if you're generous. From what Based on what we know today, okay? And I just want to emphasize it's not because we don't know that something is functional that is not functional, right? So, I mean, it's just a very hard to tell. But, we do know for for sure there's really good evidence of that still a good chunk of the genome is probably not functional, okay? Uh we can see signature of selection being really weak on some part of the genome, which suggests suggests that you could remove them and don't have any issues. So, there is uh and you know, and this part of the genome, of course, is also the part that we're interested in. It's the so-called junk DNA, right? We We're going to talk a lot more about junk DNA, but junk is a bit of a a misnomer. Um but I would insist I would I would say that at least it's better than trash. Cuz junk has a little different meaning, right? When you keep junk in your attic, you don't throw it away. You're just keeping it there. It's mostly non-functional, non-usable, may not be reused ever again, but you're still storing it. Why? Because you think, "Oh, maybe I'll do something else with it. Maybe I'll, you know, come up with a a use for it later." And that's a little bit the idea with the junk DNA concept here is that this DNA is there. It's sitting there. At this particular time, it's not doing anything, but it might be recycled, kind of repurposed for cellular function. This is actually exactly what we what we study in my lab. We study how these retroviral elements that are crashed into our genome are are buried and actually seemingly dead can be co-opted during evolution to come up with new functional novelties, right? So, that's exactly what we study. I'm going to get into this in more detail. >> Fascinating. Well, I don't think this is a totally appropriate analogy, but I sometimes I think of like a computer um you know, when I've had computers in the past, you have this small portion of the hard drive that is the operating system, the the critical functions of the computer, and as you uh you know, download programs and software and you do things on the computer and then delete things and and whatnot, you get this kind of like bloat of just extra coding that's just sitting there and it's like sometimes it causes bugs in the system, sometimes it doesn't do anything at all and you're just like, you know, over time you just get this huge bloat in the system. >> Yeah, that's an interesting analogy. It's it's and we're going to get get into it. We're going to get into it. >> Yeah, [laughter] okay, okay. >> Yeah, yeah, yeah. Um okay, so Now, another kind of shocker, I think, of the initial analysis of the human genome, and I have to say this was already reported in 2001 and it's been only confirmed since then and now we have actually a a better better estimates in fact than these numbers, but they haven't changed really that much um with the more complete human genome. Now, we know that half about half of our DNA is belong to a group of sequences that collectively we refer to as transposable elements. Those are repetitive DNA elements and this is what we study in my lab. We study transposable elements in not just the human genome. These are These elements are everywhere. Every organism has them in very variable quantities. So, there's a lot there are a lot of them in the human genome, as you can see. Um but some organisms like maize, corn, you know, uh salamanders are literally bloated with these elements. 90% of the gene of their genome of these organisms is made of transposable elements. So, really is and these genomes are actually even bigger in size than the human genome. The human genome is not particularly large overall. I told you 3 billion base pair. It seems like a lot and it is a lot. But, you know, there are organisms like some salamanders that have genomes 10 times bigger. Okay, and there is our organisms that have genomes 10 times smaller or 20 times smaller as well. So, there's a great variation in the amount of DNA that you find across the tree of life. And sure enough, the bigger the genome is, the more there are of these type of sequences, these transposable elements. So, what are they? Um so, it turns out we can think about them as sort of parasitic DNA. So, this is very much like viral DNA. In fact, as we're going to see, part of it is directly derived from viruses. These are sequences that can multiply themselves through mechanisms that have they have evolved that make them sort of selfish, we call selfish genetic elements. They have the ability to make copies of themselves. And what's that does is that whether you like it or not or whether, you know, sort of the organism likes it or not, as long as it doesn't kill the organism altogether because otherwise they would disappear. So, it's like any parasites. They can't kill their host otherwise, you know, not not as they replicate at least. So, they don't kill you immediately or rarely so. But, they would accumulate and make copies of themselves. And that's why, you know, a lot of people think that a lot of these DNA got to be maybe means nothing at all because we know that it's the result of the sort of built-in activity of these or built-in capacity of these sequences to make copies of themselves. And they would do so uh again, you know, whether you you like it or not. Okay. Uh >> So, does that mean multiple copies in the same genome? >> Yeah, so I'll I'm going to give you like a a deep in-depth portrait of those sequences in the human genome. And they are broken into like many different types. And they're part of what we call repetitive the repeat DNA because as they make copies of themselves, they they do make, you know, duplicates of themselves. But there is many different types. So you have many different flavors of these these elements in the genome. Does that make sense? >> Yeah. >> Yeah. We're going to get into these sort of like definitions in a minute. Um now I just wanted to bring up that these uh this is, you know, we we're talking about Barbara McClintock, so there she is. Uh she's a iconic American scientist, which I think is still uh not um as recognized as she should be. And she should be because, I mean, she did get win the Nobel Prize in 1983 for this discovery. Um to this day, by the way, it's the still the only woman who has won a Nobel Prize in physiology or medicine on her own, unshared. So I just wanted to bring this up. Um and she was studying um not humans. She was studying corn. In fact, at Cornell as an undergrad as a graduate student first, and then uh later on at Cold Spring Harbor in Long Island when she had her lab. And in the '40s and the '50s, so this is like really way back, she discovered this very strange phenomenon of DNA that was able to jump around and make copies of itself and hop out of the chromosome back into the chromosome. So this is what we call mobile genetic elements. So these transposable elements, of course, it's in the name. They're transposable. That's what that means. It's that they can actually mobilize in the genome. They can cut themselves out and reinsert elsewhere. And then in that process of transposition, they can also make copies of themselves. So this is how these elements can sort of invade a genome. And she discovered this in maize, and she discovered them, interestingly, through their ability, as they jump, to change, modify gene expression. We're talking about how the sequence how the genes get turned on and off. And she actually thought at the time we we at the time in the '40s, '50s, we had no idea how this works. This is prior to knowing the structure of the DNA. Okay, so this is way back. At the time, people had no idea how genes are being turned on and off, but they had pretty good understanding, and this is what McClintock was interested in, that they they good understanding that making an organism is a game of turning on and off genes. That was clear. And she thought that, actually, she had discovered the mechanism by which genes are being turned on and off, at least in corn, in this plant. Uh and in fact, she didn't call these elements transposable element at first. She called them controlling elements, cuz she thought they were controlling genes. Uh it turns out this idea was not exactly right. However, now now we're revisiting these ideas these days and thinking, in fact, it was largely right. >> [laughter] >> But what I mean by this is like it's not just it's not really the generic mechanism by which genes are being turned on and off. At least it's not the only mechanism. And it's really the important point here is like we now understand that this function, if you want, of these elements is not the raison d'être of the elements, as we say in French. Meaning, that's not really why they exist to do this. They can do this, clearly, but that's probably not why they're so successful in evolution. It's probably because this inherent selfish ability that they have to make copies of themselves. Right? Okay, so anyway, she was largely ignored at the time. People thought, "Okay, this was interesting, but, you know, it was anecdotal." Uh they thought it was a weird odd in a weird thing an oddity of maize. Um, but then later on these transposons were found in bacteria. Transposons by the way or transposable elements is the same thing. Uh, and they were discovered in bacteria, they were discovered in flies, they were discovered in yeast, and they were discovered in humans, and they were actually shown in the '80s to be um, the cause of disease in humans because as they insert into the genome they can disrupt genes. And anyway, so people starting to really pay attention and indeed it it it led her to be recognized by by the Nobel Prize in '83. And um, and since then, you know, a lot of people are more and more people are studying these elements. So, these are transposable elements. So, back to the human genome now. This is what the uh, a a more detailed picture of the human genome here. So, we retrieve, you know, our coding DNA 1%. We still have a lot of stuff about half the genome we still don't really know where it comes from. Honestly, uh, a lot of these regulatory sequences out there, but some of the regulatory sequences also map on that side of the genome, the sort of the darker side of the genome. Uh, so, you know, who knows? We don't really know. But this again, this half of the genome here, uh, the transposable elements we can recognize them. There's they have signature. You can look at their sequence and say, yes, that's a transposable element sequence because we only see this trans this type of sequence in a transposable element. That's how we disco- we we can classify those. And again, they come in many different types I'm not going to get into today. But those that I want to get into in the more in more details are those guys, the endogenous retroviruses or HERVs or human endogenous retroviruses HERVs. Uh, in total these types of sequences make up 8% of your DNA, right? My DNA, uh, anyone's DNA. So, that's a lot of DNA, right? Put that in contrast with the 1% of the the coding DNA. That's eight times more. Eight times more than your own genes are these things that look like retroviruses. So, I think it's it is fair to say that we are part human, part viruses because we also can date. And I'm going to explain a little bit how we do that. We can date when these retroviruses got assimilated, got kind of buried into our genome. And this not for some of them is not that long ago. It's you know, it's a few hundreds of thousands of years or a few million year, but it's not like doesn't go back to like the ancestor of mammals. So, these have come on, you know, they have a hitched hiked and they have come into our DNA and our lineage uh during evolution. Okay. >> So, you're saying we um literal viruses that like for example today maybe like co- if we had COVID and then suddenly at some point in history COVID became a part of our actual DNA? >> Yeah. Um exactly. And I would I'm going to explain Yes, I'm going to explain how is this possible. By the way, no one has ever described this for like COVID virus. All the almost all of the endogenous viral sequences that we have in the human genome are derived from a very specific group of of viruses called retroviruses that that HIV belong to. So, I And and I'm going to explain why that is the case. >> Okay. >> Um yeah. And that's relates to what we were talking about earlier is that retroviruses have this sort of unique ability to make a DNA copy of the RNA genome. So, here we back to like what a what a retrovirus look like such as HIV. All retroviruses kind of look like this. There's little capsules. Those are called capsid or viral particles that protects and coat. These are made of proteins that are encoded by the the viral genome themselves. And they encapsulate the genome of the virus, which as a as we mentioned earlier for retroviruses is an RNA genome. And then after they enter the cell the infected cells, they get into a cell they would uh this RNA will get copied into a DNA molecule by an enzyme that's encoded by the viral genome called a reverse transcriptase. They make a DNA double-stranded DNA copy just just like regular DNA that then gets integrated in the chromosome in the genome of the host cell that they infect. And this is this integration is not a real random process. It's also um supported by a machinery called an integrase that's also encoded by the [snorts] by the virus. So the virus has the machinery to reverse transcribe, has the machinery to integrate in the genome, and then it's going to be treated a bit like any part of our genome or like a regular gene I should say in sort of like um in disguise in a way. And it's going to be transcribed to make more copies of that viral uh genome that are going to make proteins that are going to encapsulate and get in and infect another cell. So that's the typical replication cycle of a retrovirus. All retroviruses each of them works like these. >> I see what you're saying. So you what you're saying is that viruses that have this ability to do this are called retroviruses. >> Yeah. >> And COVID is not a retrovirus. It doesn't have the ability to do this. So never has the opportunity to become part of our genome. >> Exactly. Exactly. >> Okay. >> Whereas every retrovirus needs to go through that step. Now they are you know, they are um um there has been report of different viruses, all kinds of viruses that can accidentally become even if they don't make DNA on their own, sometimes the cellular machinery can grab their RNA and make DNA copies and then integrate in the genome. But in the human genome um fixed right now 99.9% of these endogenous viral elements are retroviruses because of this, yeah. Presumably in part because of this because, you know, they they it's an inherent intrinsic, you know, ability of these viruses to get in the genome. Now, normally that should not you know become a thing in an evolutionary sense because, you know, it should not go to the next generation. However, this is when we get into the this concept of endo- endogenization. How do you become part of the genome? How do you establish kind of permanent residence in the genome of a species, right? So, um so here the idea is that if you have disintegration events of these retroviruses in s- cells that would not get passed on to the next generation, meaning we call these somatic cells. Um so HIV is a good example of that. HIV infect T cells m- most almost exclusively. T cells are somatic cells. They're part of our bodies, but they didn't they don't get passed on to the next generation. So, those are somatic integration events. So, you can also think about neurons for instance, um can be can get infected with retrovirus, there will be an integration in a neuron in the genome of that neuron and that, of course, may be may lead to some, you know, dysfunction of that cell, can lead to disease in some cases. But there is no um long-term you know, propagation of this integration in an evolutionary sense. However, and this is where it gets really, I think, fascinating. If a retrovirus has the ability to infect a germ cell meaning namely sperm or egg or their progenitors in development because, you know, we not we we we know we not when we have a slide that illustrates this, but I can just explain very briefly, but, you know, we Initially, we all born from a single cell. That's the union of a sperm and an egg. It's called a zygote. It's the first cell. And then that cell divides into two, into four, into eight, and so on, right? So, that initial ancestral embryonic cell is also, you know, a germ cell, right? Because now we don't have any gonads or anything like this. So, you know, the gonads will be derived from that cell. Everything is So, my point is if a retrovirus can infect the germ cells or their progenitors then there is an opportunity now for this integration to be passed on to the next generation from parent to offspring because the germ cells are the cells that are going to make the next generation if you do reproduce. Okay, so this is very >> sure I understand, there are there are cells that our body produces that are very much there, but they don't they don't get passed on to the next generation. And an example of that, like you said, is our neurons and T cells. Our body produces those cells, but they don't actually go to the next generation. So, any virus that infects those or becomes part of their DNA doesn't get passed on by virtue of that. >> Exactly. >> Then there are cells that that do inform what the next generation what like my child would what their DNA would look like and if it gets in that DNA then it does get passed on. >> Absolutely. >> On a very simple level. >> Yeah, no that's exactly correct. Um and and and in humans there's only two types of cells that make the next generation. It's the egg and the sperm. And they need to work together. So um to make a new organism because they bring half of the genome each. They're called gametes in genetics. Yeah, so the gametes if somehow they get infected by retrovirus that gamete can carry a retroviral insertion and be passed on. And if that gamete is used with that insertion then it can be passed on to the next generation. So you see I'm saying like it can if that happens it seems like a very unlikely thing and it is. In in principle seem very unlikely. However, we know that on evolutionary uh time frame it must have happened a lot because now we have 8% of our genome that's made of this these things and we know that the only way they could get there the only way they can get now fixed in all the individual in all humans is that they must have been integrating the germline. That's the only possible way in biology. Uh so even though it seems exceedingly rare it must have happened a bunch of times during evolution. In fact a lot more than we actually see. Probably too for us to be able to um to see now the abundance of these elements in the genome. Now things are a little more complicated meaning that once they once an a retrovirus gets integrated it can still propagate in a non-infectious way but I'm not going to get into it and multiply within a sequence like these other transposable elements that I mentioned. So in fact it can actually further um expand further expand uh in the germline. Still it's all of this has to happen in the germ line to be seen in the next generation. That is for sure. Yeah, so, you know, since we're unlikely and again, uh one thing that I want to emphasize right away and I think I'll get a little bit more detail in a in a minute is that um almost every endogenous retrovirus we have in our genome so, this 8% you know, of DNA that I mentioned earlier is actually shared with all humans across all humans. So, almost all of it, I would say 99% of it. There is some variant integration and I'm I'm going to explain that in a minute that are very recent, so they're not shared. But, the vast vast majority is absolutely shared, meaning like they come from integration events of our common ancestors of all humans, even further back, okay? So, this is this process has been going on for a long time. And, you know, collectively some people like to refer to these as the endoviron as a part like it's a sub part of the genome. It's, you know, the the viral part of the genome essentially. Uh this is the side I should like that should have just jumped to this to explain these ideas like you know, this build up of retroviral sequence in the human genome did not happen overnight. It's again the result of the the the result of an accumulation and a game that has been going on for a long time. How do we know that? Well, we know that because these days we can compare genomes. It's really amazing, but we have the human genome, of course, but we have a chimpanzee genome, we have the gorilla genome, we have the macaque genome, we have the lemur genome, we have the dog genome, we have the horse genome. We have hundreds of genome sequences available to us in the lab to study at the computer and compare, which is fascinating to do. And, it turns out mammalian genomes you know, compared to flies, for instance, or uh or plants like maize, they actually evolve relatively slowly. So, they actually pretty easy to align against each other. You can see the similarities right away. And you can really reconstruct, you know, the whole history of our species and mammals by looking at the sequences. This is called This is a whole field of science called phylogenetics. Okay. And so, when you have these different genomes, you can align them and you can clearly see that there is the vast majority, like I say, about 99 or 95% of the endogenous retroviral sequence we have in human genome, you see them in the exact same spot in the chimp genome. They're right there. Between the same exact two genes, between exactly in fact the same exact nucleotide. >> continue on, I just want to um clarify this chart here. So, the MYR stands for million years. >> Yes, that's correct. Yeah. >> And so, 90 is 90 million years ago, 75, 25, 6. And these branches So, you're saying that on the very far left, on 90, there was an an organism um >> An ancestral mammal, an ancestor, which >> Uh-huh. >> Yeah. >> And then these are are branching off into different evolutionary uh branches, is what this is showing. >> Yeah. So, what this is showing is uh a phylogenetic tree that represents the relationship of these different species. Of course, this is only a very small subset of species I put in there for this this particular figure. Yes, of course, there is a there is a branch there, of course, that I've I've cut off here, that is the common ancestor of the rodents and the primates. And that's a That was an ancestral mammal, and we know from the fossil record, all of this is extremely well documented, uh existed about a 90 to 100 million years ago, right? So, all placental mammals, so that excludes like um marsupials, uh these branched earlier, if you want. All the placental mammals have a common ancestor uh about a hundred million years ago. Yeah, I didn't show all the branches that would go to the dogs or the cow or whatever. I just to make a point that as you can see and so this what this triangle depicts here, we can map along this phylogeny along each of these ancestral lineages we can map and count the number of endogenous retroviruses that have integrated along that branch. And again, it's because this an element that we we would map here is perfectly shared between the lemur, the macaque, the chimp, the human, and all the other primate species. So, it's a it's not but it's completely missing. And you can really see this with very high precision. It's really cool to do um in the lab you at the computer can see when there's an insertion in that in all the primates that's precisely missing in all the other species. So, we can place it on that that branch, right? Um for example, elements that would map there would be present in humans but precisely missing in the chimp, in the gorilla, in the lemurs, in the macaque, and so on. So, we know most likely that these elements infiltrated the germline of humans and integrated in the in the human ancestor if they share with all humans because we can also look at that, right? And generally as as I told you, almost always these sequences are shared between all humans. Uh so, yeah. So, you can reconstruct the whole >> I just want to know it was in the last six million years. >> Exactly. Exactly. Because we know that human and chimp split between um five or six million years ago. So, we know it's got to be somewhere in that time frame, right? And again, nowadays we can do the really interesting um studies of of of different human genomes. So, if we do this and we ask, you know, of course I don't represent this in that particular tree, if we compare, you know, you and I, for instance, uh how many different human retrovirus we would find in our genomes? Well, I would have to make uh to think a little bit about this, but probably not many, maybe 10, maybe one. It's that low, right? So, it's a very few uh human retroviral insertions have occurred recently enough. And also, I need to mention right away, I have this in the next slide, we do not know, this is a really important point, we do not know in the human genome of any endogenous retrovirus that can still jump and make a new copy and become infectious. We do not know. People have looked really hard. We're still looking actually to this day in my lab and other labs, and we we have not found one. Uh that's interesting because actually in other species, like mouse, for instance, where these type of elements were first discovered, in fact, in mouse, there are some still, you know, retroviruses in their genome that are still quite active, meaning that they make virus, they're infectious, and they reinfect the germ line, and they make new copies like every generation pretty much, right? So, it's a very different level of activity than in the human genome. In the human genome, and you can appreciate this from this tree, by the way, you see that the these triangles, the size of the triangle is proportional to the number of elements. So, the big triangle here means like tens of thousands of elements. So, tens of thousands of elements in the mouse genome, we can see they're clearly missing in the rat and any other rodents. So, they have occurred very recently along that branch in the last 20 million years, and some, like I say, very, very recently. They're not shared between different strains of mice and so on. But in humans, you can see the activity of those elements are sort of plummeted, got lower and lower in evolution. That is interesting. I don't think anyone knows why. But one possibility is that we have evolved mechanisms of defense against these invaders. And this is actually, as you know, one of the things we're looking at as part of the whole project to try to find a new cure, a cure for HIV, is to actually try to borrow from or identify those defenses so that we can now re- redeploy them to uh to block HIV. Yeah, so it's, you know, there's a lot to think about here. So, that brings me to this another really, I think, really cool idea here. What can we do with this um these sequences? So, as you've seen from the the previous slide, there's this accumulation of viruses, and we can really quite I mean, you know, quite finely date them. And and I'm not going to get into it, but there's different ways of dating the insertions of these sequences than just this sort of comparative approach here. There's other sort of complementary approach. And so, you can like sort of like much like fossils, you know, like bones that you find digging the soil that you can also date with, you know, carbon dating methods or whatever in archaeology, we do we can do this sort of archaeology of the genome. Some people have coined this term of paleovirology because here what you have access is a fossil record of ancient viral infections. Very some of them extremely ancient, like millions of years old. And that is actually may sound like not much, but it's actually really cool and very important because it enables you to study the deep evolution of viruses. How were they, you know, 5 or 10 million years ago? Are they look Do they look like the modern viruses? So, that informed us a lot about how viruses change, how viruses evolve by by looking at this fossil record. Um now, I >> want to clarify >> Yeah. >> So when these viruses become part of our genome it's not like we have 500 active viruses that are battling their way in our in our bodies, but they're actually not not awake or not not I don't know what the word is. >> Yeah. >> For the most part? >> Um >> active? >> So it depends if you're asking when. So, when they first got into the genome, right? Then they made that integration they must have come from an active progenitor. So, at that time that's the thing that's I didn't emphasize this enough earlier. It's a very dangerous game. Truly to get one of those genome in the germ line in the germ cell and pass it on to the next generation because you can imagine from that one cell you're going to make all the cells in the body. They all derive from that single cell. So, that means that every cell in the body essentially will have that integration of that retrovirus. Technically half of the cell, but I'm not going to get into the genetic details, but still it that means that the potential the the potential for these to infect uh to to be infectious and to destroy the organism is is huge. It's catastrophic. It's much worse than you know if you get a virus in a single T cell. Here you get all the cells that carry these. And so, if this virus has the ability to make more viruses it would be it it seems like it would be like really catastrophic. So, it's possible that some of these elements are sort of dead on arrival. So, they that's what enables them to move on to the next generation because in fact they cannot replicate. So, that's a possibility. It's hard to to answer this question. Uh but we know from studies in animals that do have actually active germ line infections and endogenization going on. We know that's not the case. There's viruses that make it to the next generation that are capable of making more viruses. So, the best studies of these were in mice, but also interestingly, um it's been seen it's been shown that in koalas, the cute, you know, um marsupials from Australia, the koalas have an ongoing invasion of endogenous retrovirus happening right now in their genome in live. To the point that there's population of the koala of koalas in Australia that are completely devoid of these these elements in their genome. And then there's another population I can't remember if it's the north from the south the south of the continent. They They have They are invaded. They have hundreds of copies of these things in the genome and they clearly still active. And the koala appear to sustain this. They do cause disease. So, we're going to talk about the consequence of these these koalas have lymphomas that are associated with these and then kill them. So, it's not without a cost, okay? But some survive and some make it to the next generation and they fix a new retrovirus, right? So, there's an ongoing invasion of the koalas. So, there's a lot to learn from studying this because in humans, we don't see this. As I said uh early earlier, no one has seen a new jump, you know, so you look at a uh a baby and you ask, "Oh, does this baby has an insertion that I didn't see in any of the parents?" This is called a de novo integration event. We don't see this for any endogenous retrovirus. No one has ever reported this. Again, really would be a different picture if you look at the mice or koala, you would find such new integrations. Yeah. >> And so, um just going off of that a little bit. So, if if a virus has the ability to become part of the genome and become endogenous, um this would have to happen like it couldn't just happen on an individual level because for evolution that one individual isn't going to create all the offspring. It there's a whole population like you said of the koalas. So that then this has to happen in the virus there has to be this change that happens that's that's transmitted to all these koalas at the same time that all the virus has the ability in all of them to go in the genome? >> Yeah, so no but and that's that these are great questions. So no I mean the answer the quick answer is like once this insertion occurred in one individual it would it can only be passed on from parent to offspring actually. So actually I know that's why it seems like incredibly uh unlikely that we would get to what we call fixation. So fixation would be like when all the individuals in the population have this insertion. [clears throat] It means they all come from the same ancestor that had this insertion. So that's the Wow. You point yeah. I know because That just makes it even that more amazing that that that's the case. Absolutely because because the the population genetics is a whole field of that that study this this process. Population genetics tells you that you know the likelihood of fixation of any given insertion is exceedingly low. In fact it can can be calculated. Is the implication of that then that there were other other ancestors and other trees that all died off that had their own endogenous retroviruses that never that aren't here because those trees all died for whatever reason? Exactly. Wow. In fact the vast majority we don't see. In fact yeah the population genetics tells you the math tells you that for in humans uh the probability of an insertion to get eventually fixed like this meaning like only individuals have it and so on, it's 1 divided by the effective population size, which is 10,000. So, it's exceedingly low probability. And yet, as happened in the Yeah, so what we see in our genome today absolutely the tip of the iceberg of what has happened, of course, across large evolutionary timescale. But, you know, that's the thing in the koala, right? Ongoing right now. And you know, most of these insertions, they are not shared between individuals, right? Because they're very recent. So, they haven't been shared, but some I don't know to what extent they are shared. I don't actually know. I don't study the koala, so I don't know. But, yeah. All right. So, to illustrate this idea of paleovirology, I put one slide from an old paper of ours. But, it relates to HIV, so I thought it would be interesting to to bring it up today. Um but, I don't know. This was published in 2008, so some time ago now. We stumbled on something really interesting. So, in the human genome, there are no endogenous retroviruses that are directly related to HIV. Okay? So, there's no endogenous HIV. Meaning an HIV that gets passed on to the next generation as part of the chromosomes like this has never been described. We know that HIV can be transmitted, you know, from uh infected mom to a baby, but that's a different route of transmission. That's an infection. That's a horizontal transmission. It's not a vertical transmission. Okay? But, um and so >> So, is that a misnomer then when we say vertical transmission? >> It should be, yeah. It It is not correct. I know it's being used, that term is misleading. Yeah. >> And that's super confusing because when we think about HIV HIV, we it it seems like it's something that we passed on genetically, but by virtue of the fact that we can the mother can become undetectable and therefore have a negative child proves that it's not. >> Yeah, exactly. No, no, it is not a genetic transmission. So, it should not be called vertical transmission. If you see these people say that say that, we should we should correct this. It's not correct. But, it's often referred from like mom to offspring, right? So, that's kind of like the same it's still a parent to offspring transmission, right? But, it's not a genetic one. It's an infectious one. >> Okay. >> Exactly. >> That's a good distinction. >> Yeah, yeah. Um Yeah, so, you know, us and others, we've been sifting the human genome for like things that look like HIV. We'll be interesting to find a a fossil of HIV. First, it would tell us where how how long we've seen genomes like this. So, you know, HIV is part of a a group of retroviruses, a subgroup of retroviruses called lentiviruses. And lentiviruses certainly exist in other species than humans, right? And we know that you know, most likely we did acquire HIV from chimpanzee so-called SIVs, simian immunodeficiency virus. Uh we also know of viruses like this circulating in gorillas as well. Um and and it's you know, and horses are are are infected with a virus lentiviruses. So, lentiviruses are kind of widespread. But, you know, if you look at their sequence of these other these lentiviruses from different species, the the infectious ones, you arrive at an origin and try to infer for how long they've been around. You you get a you get a date of of birth, if you want, for this clade on around a a few thousand years. You know, that's kind of what the the the sequence of the infectious viruses tell you. So, we're really interesting to see can we find, you know, old lentiviruses buried in in in genomes. And in fact, um before us, there's a group in the UK that found an element an an endogenous lentivirus in a rabbit in the genome that was fixed, you know, meaning like all the rabbits had these insertions and so on. And this was clearly a relative of HIV. You can take the sequence and make it into a a tree of viruses and it would go and group with HIV. So, it was the first description of an endogenous lentivirus. And us and others >> uh explain the chart here a little bit for folks. So, it's a on the the left it says time in million years ago. So, we're going for on the bottom that's today and then going back in time. And then we're seeing four different variations of this virus. And so, on the the top line that's the most ancient variation and then you can see as you go down these changes, which I'm sure you'll >> Yeah. Yeah, I'm sorry. I'm just putting this this picture without really attempting to really explain it. I was just more of an old illustration. It's an old slide I have in my slide decks. Um but yeah, this is sort of depicting the sort of structural evolution of this type of retroviruses. You know, this is what's shared with all retroviruses. This sort of the canonical structure of all the retroviruses like literally all of them. They look like this. Then when you get into the lentiviruses, you know, you see additional things come up that they evolve new genes. And this sort of there's a picture that emerged that was this picture was put together to illustrate the sort of increasing kind of complexity of these retroviruses. Um and again, the reason why we can reconstruct this these kind of evolutionary scenarios is because now we uh we and others have found fossils of the lentiviruses buried in genome that we can date using different methods. So, you know, this one so the the one in in rabbit was dated to be at least 12 million years old. So, that's really old. And that's because it was shared with like hair and other other like rabbit-like creatures at the same spot. So, it's and it of course nowadays it's got we know it's representing on this graph as being like sort of intact, but in fact they are like really fragments that are found in these genomes that are kind of put together just like the bones, you know, here. They're found in different places in the genome, but you can reconstruct the entire skeleton of these ancient retroviruses by looking at these pieces. And and I could say us and others, we were actually excited to find that in lemurs. So, these are primates. They're not, you know, they're not too close to humans, but they're not as far as rabbits are. They are in in you know, they are in the in a bit the island of Madagascar. And in these lemurs in the genome sequence, we found again, bits and pieces, right? They don't want to emphasize they don't look like perfectly clean like this, but you find bits and pieces, you can put them together kind of like a puzzle or kind of like a skeleton like this. And and reconstruct what the retroviral genome looked like. And then it you can compare them to, you know, the ancient one and also you can compare them to like a modern, you know, current HIV or SIV retrovirus. So, I'm bringing this up because this two really interesting point here is like first, I mean, you can get information, knowledge about the deep evolution of viruses, so that's really interesting, important to understand you know, their basic like biology. But also unlike unlike the fossils like this that are made of bones, you know, at best you can reconstruct the skeleton and display it in the museum, but here you can actually potentially put them back together as functional viruses. And I know this might seem like a very bad idea, a very not good idea, but it can be done in controlled condition, it can be done in with like safety device so that it can't escape from the lab to kind of study how do they replicate? And also you can use them to like compare their resistance to drugs and all of all the really important translational studies. Um so, I wanted to bring this up because it's kind of like sounds a little bit like science fiction, but it here is actually a hard science, right? You can in fact put back together uh, ancient extinct viruses and, you know, study how they propagate, how they replicate. And now, I'm going to return briefly and I'm going to finish with this, in fact, uh, about what we know about human endogenous retroviruses, or at least a brief summary of what we know about the human genome, what do we have in there. So, in total, I told you 8% of our DNA, so it's that's a lot of sequence to look at to look at. Um, and in fact, uh, this 8% is made up of 400 pieces. I wrote pieces because they're like fragments, segments of DNA uh, that are definitely of homology, a sequence similarity, you know, to endogenous retroviruses. So, that's a lot of things to look at. Again, you know, to give you an idea, we have about 20,000 genes that encode cellular proteins, but we have 400,000 HERV segments in our genome. So, it's a lot It's a Yeah, it's it's really a lot of work to do to look at all of them, and um, certainly we know them much less than genes. So, it's still very much like, you know, unknown stuff, right? So, we're a bit of again like the dark dark matter of the genome, the dark corners of the genome. I would say, you know, this is still very much understudied, um, in the field in genetics. Um, but there's a growing number of scientists that are now turning to these parts of the genome and looking at them. So, we, um, you can classify those into different families by their sequence relationships. You can make all these trees with one another, and you come up When you do this, uh, you can slice it in different ways, and you end up with hundreds of very different retroviral sequences, meaning they they they they they derive from very different types of retroviruses. Things that are just as different as HIV and say another retrovirus that's very well studied is called murine leukemia virus. That's an endogenous and exogenous one, by the way, in the mouse. These are These are very different, but they they still have the same exact organization, you know, like the organization of retrovirus really never changes. It's like these long terminal repeats, these are like regulatory sequence that enable the viral genes to turn on or off in cells. And then the genes that encode the proteins that are important for the replication of these retroviruses are there. But of course, when you go through the human genome and look at all these pieces, often what you have is only part of these, right? You only have part of the LTR or part of that sequence, part of that sequence. And again, you know, you can kind of reconstruct the pieces and look look like look at what the ancestral retrovirus looked like, but I want to emphasize that a lot of what we see is clearly like decayed material, right? Material that you can tell shouldn't be able to make a virus again, okay? It's really disrupted with a lot of mutations and changes. But anyway, I want to emphasize that there is a great diversity of retroviruses buried in the genome. But again, none of them are closely related to HIV in the human genome. And us and others have been studying these because, you know, we're being told initially, you know, you're being told, "Oh, this is just dormant. It's just sitting there. It's not doing anything." But, you know, the more we look, the more we see that those sequences actually can have important impact, functional impact, on the on the on the organism, including humans, right? And I'll give you just a few examples in the next slides about what these could be. Um but one of the things that, you know, we and others have been studying is how do they respond to change in the environment, including infection. It turns out that it was discovered more than 20 years ago by Doug Nixon's group, actually, and others, uh who study, in fact, uh people living with HIV, in their in the T cells, you do uh observe that certain endogenous retroviruses are activated compared to T cells that would not be infected. And you can recapitulate this in vitro. You can take T cells that are not infected, infect them with HIV in the lab, and you will find that a subset of endogenous retroviruses get activated at the transcriptional level. They start making RNA, which to me is >> Woah. >> Yeah, it's a sort of a fascinating interaction that is going on because we're talking about, you know, distant cousins of HIV. How come they kind of respond to the presence of HIV when when when HIV enters cells? We have a pretty good idea of why that is, by the way, yeah. It's because they probably So, when HIV infect a human cell, it will turn on, you know, it will be recognized as a virus by part of the immune system of that cell. And as a result, there's a whole cascade of things that happen in the cells, and that part of this cascade of things is to turn on genes that control infection, or aim to control infection, attempt to control infection, and to respond to the infection. And the same mechanisms, or the same factors that control those genes, also control these endogenous retroviruses. >> Mhm. >> Uh >> And so, are these endogenous retroviruses that are impacted by HIV infection, are they complete, or are they fragments that somehow get turned on? >> Yeah. It's an excellent question. Um we're very much looking at these in a lot of detail. But as I told you, like so far, no one has ever found an endogenous retrovirus that would be capable of being fully infectious and replication competent. So, you can look at them and some are almost complete, right? They look like kind of like almost perfect, but not perfect. You can see they have a one or few mutations that are predicted to disrupt the activity to make a virus. So, we do not think at this point in time that any of these retroviruses can kind of wake up and make more virus. But but we know that some can make partial can make parts of a virus. And actually, you know, Raif, in the lab we're actually studying these right now. Uh really uh these days we're wondering if some of these retroviruses can actually mix and match with HIV >> Mhm. >> as the cells get infected and make sort of chimeric kind of viruses. In fact, something like this has been previously reported in the literature. Uh so, we're very much interested in this question. >> And when you say um that they can some of them can make fragments, is that similar to uh like for example, myself, even though I'm on effective treatment and I'm undetectable, my latent reservoir HIV is still creating fragments that are essentially inert, but they still cause an inflammatory response in my body, which leads to comorbidities and all that. Is that similar that some of these incomplete endogenous are they're creating fragments? >> Yes, absolutely. And in fact, this is an area of very intense uh investigation at the moment in the field because you can see how this can, you know, exacerbate inflammatory response if the same molecules, even partial, can trigger an inflammatory response, it would imply that it would implicate that these endogenous retroviruses can kind of uh exacerbate, you know, >> Yeah. >> uh make it worse. Yes, and this is being studied and and and yeah. So, and this is taking me like to the to next slide and I I'm I promise I'm almost done with this. Is like um the impact, right? Is are these things at the end are they good? Are they bad for us? Uh probably both is true. And I'm just going to give you a a few vignettes, few examples to illustrate this. Now, I also wanted to mention before I move on that this is important for what we're doing for the the HIV cure. Uh it is known that the vast majority of those, when I say that some wake up, but they're a small minority, right? The vast majority are really clearly deeply dormant. Um you know, I don't want to say they're like um Sleeping Beauty uh because you know, we don't want to think that they could ever wake up. But they look like they're really dead. Or at least they they are really tightly really tightly controlled and maintain in a deep silent state, a deep sleep. So, this is what's of interest of course to us because and people or the collaborators in the Hope Collaboratory that maybe there's something we can learn or how our cells, our somatic cells and in particular T cells cope with that, right? Cope with this mass of elements which are largely dormant. Again, not all of them are. Yep. So, this is just my one summary slide about how to think about the good and bad, the evil the good and evil of endogenous retroviruses. When you know, there's no question that when these sequences were first discovered, in fact, they were first discovered in uh in mice. And how they were discovered? They were discovered because they caused cancer. So, in mice, endogenous retroviruses integrate and they they directly cause cancer. Very well documented. In fact, when they were then later on later on seen in the human genome in large numbers, people immediately, many cancer biologists focused on those sequences thinking they would have found, you know, the cause of cancer essentially in humans, like it was, at least for some cancers in the mice. But, they came back empty-handed after many very you know, decades of research actually. There is no evidence that they cause cancer the way they do in mice or in koalas as I mentioned, you know, uh or in other organisms. And the reason is pretty simple. Like because they cannot make new viruses and new integration in the genome as far as we can tell, they know they're not replication competent, they're no longer mutagenic, and that is how they cause cancer in mice largely is because they insert into genes that suppress cancer or they turn on genes accidentally by inserting next to them that should not be turned on that are called oncogenes that that cause cancer. So, that's what's happening in mice, and it's not really happening in humans. However, people have looked at this and, you know, as we just talked about, they do make partial viral products, right? So, they're not complete, but they still they can encode stuff. They can have also regulatory activities, and it's been seen that they can actually have sort of oncogenic activities when they wake up in the wrong place in the wrong time, you know, if you want. So, if they disregulated, it's easy to see how they can turn on and off adjacent genes in the genome that they normally reside. So, they're not a new insertion right in the genome, but they are disregulated, and now instead of being silenced, they wake up, and they may activate flanking genes. So, there is a few examples of these where clearly the endogenous retrovirus reactivation is responsible for the apparent expression of oncogenes. There's a couple of examples, not many. Um this is still, you know, I'm not going to spend too much time, but this is really an active area of research because there's many disease states and human diseases where a subset, again, it's sometimes very specific endogenous retroviruses have have been have been shown to kind of reactivate like this or being making RNA, making proteins. Again, they cannot make a full-blown virus as far as we know, but they can make products. Whether these products that are overexpressed in these conditions are part of the disease is really up for debate and we there's not a lot of evidence at this point of that, but people are really looking at this because it's been seen, you know, in neurodegenerative disease such as ALS or Alzheimer's disease that a subset of these elements get upregulated in the brain. Do they contribute to the neurodegeneration? That's an open question actively being uh studied right now. Uh and I mentioned this earlier that indeed in condition of viral infections, including HIV, we do see that some of these products get reactivated, including in T cells. So, that could be like sort of the the dark side, you know, really of these elements is that they may really contribute uh to human diseases. Now, in my lab, we have actually been focusing on the brighter side of things um because that's life, right? It's never like black or white, it's always sort of in between. Um but we've been looking at the at the at the more constructive activities of these sequences in normal human development or human physiology. Um I need to say, you know, so this is kind of the idea of uh repurposing. You know, I mentioned this earlier, the idea of junk. Well, junk is not trash. Junk may be recycled, right? And it and we know and we've been studying this process in not just the human genome and and and other people have been studying this process. Um it's pretty common, actually, it turns out. We don't know how common it is, but a lot of these sequences have been co-opted to make new things, new good things, you know, cellular to to promote cellular innovation, genetic innovations. You know, so this is a very general question in biology. Where do new things come from, right? Obviously, you know, even though, you know, I told you like human and chimpanzee are very similar genetically, and we're we're not the same genetically, right? There's clearly reasons, genetic reasons for why they we look different. And we do and we think differently, and we can, you know, one species has language, the other one doesn't, you know? So, they are really genetic novelties that have come along the way that make all the species different. And we think that the co-option, the recycling of endogenous retroviruses and other types of transposable elements like these um has really fueled, you know, genetic innovation. And I think it's not really, you know, challenging to understand these because they're complex sequences. They come with their own genes. You know, they were not there uh to make new genes, but they nonetheless sit in the genome, and as I as we explained, as we as we discussed, they are uh to some extent active. They make RNA, they make proteins. And so, once in a while, occasionally, at least, some of these products may become beneficial for the species that harbor that express them, or the individuals that express them. >> It's so interesting, and you went you totally went where I was going to go with this, and and what I was going to ask you about is that if there is situations where it actually ends up being beneficial, and I would think in nature, when you have a working system and you throw something random in it, the probability is that it's going to cause harm versus that it's going to support it. Um just by by design. Um >> Or or do nothing at all. >> like >> I'd say that's maybe >> Yeah, or do nothing at all, exactly. >> It depends on the organism, actually. Yeah. >> But it sounds like you're saying that the potential for it to be beneficial is a little bit more than just what it would be by chance. >> tell because we don't Yeah, it's Yeah, it's it's It is kind of hard to tell because we don't have the the we found the the the catalog of, you know, in the human genome for instance, the catalog of transposable elements and endogenous retroviruses that are really have turned into like functional and important things. We really have a very uh partial catalog of these. So, I you know, I can't really uh speculate on the frequency of these events. But it's really a different process when you get a retrovirus in your genome than when you get a point mutation, you know, which is a change of a single base pair. You can see that how the change of a single nucleotide certainly sometimes can have a big impact. But most of the time it's going to be a minute impact on the function of the genome. And most of the time no impact at all. This is actually known. Uh but when you get an integration of a complex thing like a retrovirus that comes with its own regulatory sequence, its own genes, all the bells and whistles that we know, unfortunately, about HIV that defeats that makes HIV defeats us, uh come all like in one package, right? So, I think it's, you know, the at least I would argue that the potential for these to introduce um you know, dramatic change is here. And again, it can be for better or worse. Here is an example. Uh that is pretty probably the most spectacular example that we know so far of a retroviral sequence that has now become essential for all of us to be born. Okay, at least that's Carl Zimmer, you know, he's a famous science writer, fantastic I think science writer, wrote this article uh for the New York Times that summarized this, but there's many other articles you can find online on that story. It's a story of a protein called syncytin. And this is a, you know, it looks like a normal gene now today in the human genome. It's annotated as a gene that encodes a protein called syncytin. Uh and it turns out that this protein is entirely derived from an ancient retrovirus. And in fact, it's derived from the part of retrovirus that makes the protein called the envelope protein. And you may have heard that the envelope protein, by the way, is the same thing as the spike protein of the SARS-CoV-2 that we all got, you know, injected. It's an envelope. The envelope is the part of the virus that enables the virus to enter a new cell. It's a part that recognize a receptor on the cell and that enables the fusion of the particle of the virus with the host cell. That's what an envelope does for a virus. And amazingly, in a sort of crazy evolutionary story here, that ability of the envelopes to fuse cells has been repurposed for the making of placenta. Placenta is the defining organ of placental mammals. You know, all mammals have a placenta. You know, all placental mammals. Um and the placenta, there's a layer of cells around the placenta called the syncytiotrophoblast. This is a a layer of cells that are fused to one another to make a whole big bag of cells, and that is the direct interface. This is a embryonic tissue, actually, right? The placenta is derived from the embryo, not from the mom, but it fuses it invades the maternal tissue to enable the implantation and of course the feeding of the developing embryo. Turns out that layer of cells and the fusion of those cells is accomplished by this protein syncytin. And that kind of makes complete sense because that's what envelopes do. They fuse the viral membrane with the new the cell the membrane of the cell that they infect. And that's the same process that has been co-opted to make that layer of cells around that defines the placenta, actually. So, that's a really spectacular example. I think everyone would agree. Um because we were all born. We all need a placenta to be born. And we owe this this process to at least in part to the co-option of an endogenous retroviral envelope some time ago in evolution. The craziest part of the story, I'm not going to even get there, is that different mammals have different syncytin that are derived from different retroviruses. So, this co-option has happened over and over during mammalian evolution. Different retroviruses have been borrowed in different lineages to kind of tweak a little bit the placenta and to make the process of fusion a little bit different. So, it's absolutely fascinating developmental biology. But, I think it's a clear example of a protein of retroviral origin that's now essential for um you know, the development of all humans and actually many other mammals as well. >> So, is the ability to have um to give birth not via an egg, but through the placenta, is that a defining characteristic of a mammal? >> It is a defining characteristic of a placental mammal. A marsupial is that's why they have the pouch and the baby in the pouch because they finish development in enables also the placenta enables an extended pregnancy. And you can keep the baby the embryo until it's, you know, really fully developed. Yes. >> So, that is what what you're saying with >> is not how it looks like, by the way. This This is a vision of the placenta prior to like knowing, right? So, don't don't don't get this picture. This is not how it looks. >> [laughter] >> But, what what you're saying is that as the result of this retro virus from a long time ago that became part of our genome, we were able to form the placenta. >> At least part of it, yes. The essential a very essential part of it. And I know maybe some of the viewers might be wondering what what happened before that happened then. You know, what was the protein? So, that's a really good question. No one has found a syncytin-like protein that would be shared with all placental mammals, meaning all the mammals that make a placenta. One would argue there must have been one. Possible. It's It must have been some some mechanism to do this because all the placenta share, you know, the cell of syncytiotrophoblast. But, what it looks like is like there were innovations sort of anatomical innovations in the placenta that was fueled by the acquisition of new envelope proteins. And in fact, you know, I I mentioned syncytin in human genome, there are two actually. There's syncytin one and two. You know, I always simplify things here. But, there are two proteins derived from two different retroviruses that were captured at different time points in evolution. They both primate-specific, but one is older than the other. We and others have studied this in a lot of detail their evolution. Uh there are two, and in the mouse there are two as well. But, they are different. They come from different retroviruses. So, there's been a sort of like We don't understand why fully. But, there's been a uh sort of a turnover or revolving door of these proteins getting co-opted to make this apparently the syncytiotrophoblast. This is very well characterized in the in the mouse, by the way, cuz in mouse there's two syncytin, A and B. And there's a group in France that made knockout genetic knockout, so they removed the gene altogether. And of course, the mice couldn't develop the placenta. So, we we know they are really clearly essential, at least in mouse it's very well documented. Okay. Um I'm going to wrap it up and give you like one one example of like a constructive uh constructive retroviruses. And actually, this example is really what we've been studying in my lab for the last 10 years or so. We've been looking for cases where endogenous retroviral sequences have been co-opted to boost human immunity. To actually reinforce or drive the immune system. Because, you know, how cool would that be if retroviruses were now being used, you know, to fight against viruses, right? So, that's the idea of fighting fire with fire. And we found a few really cool examples of that. And I'm going to I'm not going to show you all any data or anything. Uh the So, most of this work is already published. Um so, people can look at all the details in the papers, but, you know, working with a really amazing post-doc in the lab, we were able to show that a bunch of immune genes, so those are genes that encode like antiviral proteins, antimicrobial proteins, including proteins that actually are um fighting HIV, actually. These genes, these are like really important genes in immunity, are in fact regulated by the regulatory sequence that once regulated the expression of retroviruses. Okay. So, these long terminal repeats, these LTR sequence, which are like dispersed in hundreds of thousands of copies in the human genome, a subset of them are now being used to regulate our own genes. So, you see, I like this because it goes all the way back to Barbara McClintock. I told you she thought that this is how genes were regulated. And she was This idea was completely dismissed. She was basically ridiculed. But, it turns out now, we and others have found really clear examples of this, including in humans, where our own genes are regulated by the by these rogue elements that are clearly now, you know, being co-opted, right? They've been like repurposed for this for this purpose. So, uh so that was an interesting example. And then there are other examples out there of the proteins that were once encoded by retroviruses. I mentioned the envelope proteins and another example would be these capsid proteins that are now during evolution turn into antiviral proteins. How cool is that, right? You have once they were serving a virus and now they're defeating a virus. So, these are the things that we were really interested to look at a few years ago. Uh this had been really well described actually in the mouse both for the capsid and the envelope. It turns out it was known that some of these envelope that are uh in the genome encoding in the genome of mice can block against the infection of incoming viruses in the mice. And a couple years ago we had a paper driven by amazing grad student in the lab, John Frank, uh on a protein called suppressing, which is encoded in the placenta. So, we return to the placenta, by the way. This is kind of a hot spot of retroviral co-option. In the placenta this protein that we all make, I mean at least, you know, in a developing embryo, uh this is again this is not a maternal protein, right? The placenta is derived from the embryo itself. So, we all as embryo express suppressing regardless of our or or sex or or biological sex, okay? Just want to clarify this. This is a embryonic protein. Suppressing we found we reported in this study blocks can block at least in cell culture. Need to say this were all experiments done in cell culture, not in vivo, of course. In cell culture this protein can block against infection by a large group of retroviruses called the type D retroviruses. These are gamma retroviruses. They're not related to HIV. They are different retroviruses. They are known to infect a bunch of species, mammals, uh including bats, including primates. I'm infected by these type D retroviruses. Humans can get infected, but there's no report of you know, infectious type D retrovirus circulating in humans. We think that perhaps in part it's because we have antiviral proteins at such as suppressing that help us fight against these retroviruses. Of course, this is you know, this this is a speculation, but we clearly show that this suppressing can block against entry infection by type D retroviruses in vitro. So, uh and by the way, we think it's really like probably the tip of the iceberg of what this what what can what these proteins can do because we found that they were about a thousand different envelope proteins that were potentially expressed in human tissues in different tissues, not just in the placenta, in the brain, and elsewhere. So, we think that could be um you know, they could be a pool, a reservoir of potential antiviral proteins encoded in our own genome that we don't even know about. All right, I'm going to stop there. Uh but you know, this is kind of the idea that the opportunity for defeating the enemy is provided by the enemy itself, right? So, how can we learn from all this business um new ways to combat HIV? And this is part of this whole collaboratory that's NIH funded that you sure you talked about before that you involved with, right? Um HIV obstruction by program epigenetic is is is the the big the complicated word for it. Um yeah, and as you know, this has involved many institutions, an international project, many associations. It's really fantastic to be to be part of this for us. This is the most translational research we've ever done in my lab. And really uh I I really just really enjoy it. I think it's just very inspiring to work with this this team, including you. Um yeah, so we hope that it would make a difference. And so you know the strategy is that of block, lock, and stop. Meaning that we want to develop new ways of repressing preventing, you know, what you described earlier, preventing the reactivation of this latent HIV that sort of hiding in the genome even when you are being treated with antiretroviral therapy. The potential is almost always there to have a reactivation of of one of those latent retrovirus. So one way to prevent this would be to put them in deep sleep essentially so that they cannot no longer reactivate. And so that's that lock phase, that ultimate stage where, you know, we would permanently silence the HIV provirus. And so of course here's for us the idea here is that can we look at how endogenous retroviruses are being silenced in T cells and learn how it's done and and then redeploy reuse the same strategy to design repressors of HIV. Sort of you know, Melanie Ott, the you know, lead PI on this project likes to think about can we accelerate evolution in a way and, you know, sort of endogenizing the sense in a sort of conceptual sense HIV. It's different of course because it wouldn't go through the germ line, right? That's a different it would only be in T cells. The intervention would be only in T cells. And yeah, and so here the idea for us for our group is to try to borrow from nature. And it turns out in another like sort of fascinating evolutionary twist to the stories that I mentioned today our genome encodes a battery. I like to think about this as an army of proteins, cellular proteins, that are dedicated to silence these endogenous retroviruses in a very tight way. And this has only been discovered in the last 10 years or so. So, it's a pretty new discovery. It's kind of amazing because there are 400 genes in the human genomes that encode this type of proteins that are called KRAB zinc finger proteins. And uh again, these are not retroviral proteins, right? These are just cellular proteins. And until recently, we really had no idea what they might be doing. But it's become clear now that each of these genes encode a different KRAB zinc finger protein that can bind, physically bind onto the DNA of specific types of transposable elements and endogenous retrovirus. So, it's almost like for every type of family of endogenous retrovirus or transposable element families, for every one of those, you have a matching KRAB zinc finger that binds to it and block it, silence it. And the machinery by which they do this is actually pretty well understood. Um however, we still very few studies have been conducted of these type of proteins in T cells. So, we're studying now those KRAB zinc finger proteins in T cells. And basically, the question is like, well, are they really doing this in T cells? Cuz we're not sure about that. So, we're looking at these. And how are they doing it in more detail at the molecular level. Yeah. And so, I have two um great lab members involved in this project, where we want to try Sabrina and Weihu, where we don't try to harness these KRAB zinc finger proteins to block HIV. So, I won't get into uh you know, what what all the things we've done in this area, but give you just like the the the goals here. So, the goals once again is to kind of map which are the KRAB-zinc finger proteins that are binding and silencing HERVs in T cells. Dissect the mechanisms of silencing. How does this work in a molecular way? And then also what we're trying to do and we're excited because we actually have identified some KRAB-zinc finger proteins that actually can directly silence HIV. Turns out that they can actually bind HIV and silence HIV. We're in the in the process of validating this at the moment, but we're pretty excited about that. This was a bit of a surprise, actually. Um and then we hope to use this knowledge to design now new repressors of HIV. Can we borrow parts from these KRAB-zinc finger proteins and make new repressor to silence HIV in a permanent way? I hope that that that does it. >> [laughter] >> Yeah, that's a fantastic um >> That's a long That's going to be a long episode. >> We covered a lot. >> Anyway, always happy to do it again and you know, delve delve into any one of those topics again. You know, I I love doing this. Outreach is >> I think it's it's fascinating. Well, Cedric, that was probably fascinating um discussion or presentation on your part. So many So many interesting threads to pull, so many questions and and uh just different branches off we could go and I and I'm I'm looking forward to talking about KRAB-zinc finger proteins with you cuz this is specifically related to HIV, but I wanted to give everyone that foundational understanding of HERVs and HERVs. Um Cedric, thank you so much for taking this time uh to explain this concept in in in a way that hopefully uh people can understand. By the way, guys, if you didn't understand everything, don't worry. As long as you're picking up little little things here and there, um over time, hopefully you can add to your lexicon and get a better understanding as we repeat certain things over and over. Everyone watching, please comment below your thoughts, your questions. Were you able to follow along? Did you learn something new? Would you like to see more content like this? Any specific questions for Dr. Feschotte? I'd be happy to follow up. Please like this video, subscribe, hit that notification bell, and share this with anyone who might find value in this content. These are the best ways that you can help support me and my channel. Until next time. Cheers.