Submind YouTube summaries
Thumbnail for Course Summary - (Bioinformatics S15E1)

Course Summary - (Bioinformatics S15E1)

Watch on YouTube

Video summary

Bioinformatics is defined as an interdisciplinary field that leverages computer science tools to address complex biological questions, distinguishing between raw data and actionable knowledge while operating in *in vivo*, *in vitro*, and *in silico* contexts. At the heart of this discipline are DNA sequences, which serve as the primary entry point for studies ranging from phylogenetic tracking of viruses like the coronavirus to whole genome shotgun sequencing. The curriculum explores various applications of DNA technology, including diagnostics, forensics, and biotechnology, while detailing specific sequencing methods such as Sanger sequencing and Next-Generation Sequencing (NGS). Students learn to interpret sequencing outputs, manage large computational files like FASTQ and BAM, and navigate the NGS workflow from sample preparation through alignment and variant calling. Furthermore, the course clarifies fundamental definitions, noting that a "gene" can refer to either a unit of inheritance or a complex genomic sequence containing introns and exons, and introduces transposable elements discovered by Barbara McClintock alongside regulatory mechanisms involving enhancers and insulators. Beyond DNA, the course expands into RNA and protein biology, covering diverse RNA types such as mRNA, tRNA, rRNA, and regulatory RNAs like miRNA and lncRNA that play crucial roles in gene expression and splicing. The RNA-seq workflow is examined with a focus on measuring expression levels rather than identifying SNPs, utilizing tools to predict secondary structures based on free energy. Similarly, protein studies distinguish between apoproteins and holo-proteins, explaining how amino acids form chiral structures that progress from primary sequences to complex quaternary arrangements involving multiple polypeptides. Advanced topics include the use of machine learning in structure prediction via AlphaFold, separation techniques like 2D gel electrophoresis, and the evolutionary relationships between orthologs, paralogs, and xenologs resulting from gene duplication and speciation events. The study of metabolites is also introduced, classifying them as primary or secondary compounds and outlining mass spectrometry processes involving chromatography and fragmentation analysis. Genetic analysis techniques are further explored through the examination of inheritance patterns, where traits are categorized as either Mendelian or complex polygenic characteristics influenced by additive mixing and dominance. The course explains how linkage determines the co-inheritance of close genes on a chromosome while recombination separates distant ones, utilizing two-point and three-point crosses for genetic mapping. To study these effects without the complications of heterozygosity, recombinant inbred lines are created through brother-sister mating to produce genetically identical populations, whereas back crosses lead to genomic imbalances that fail to capture dominance or additivity. Consequently, F2 crosses are highlighted as a superior method for investigating both additive and dominant genetic effects, despite their limitations regarding large associated regions. These concepts bridge into modern mapping strategies, comparing QTL mapping in structured populations with Genome-Wide Association Studies (GWAS) used in outbred human populations, where results are visualized using Manhattan plots to identify genomic regions controlling specific phenotypes. The final segments of the course address the statistical rigor required in genetic analysis, emphasizing the importance of distinguishing between effect size and likelihood when comparing homozygote groups. Since researchers often conduct numerous statistical tests simultaneously, multiple testing corrections are essential to avoid Type I and Type II errors, such as misidentifying outliers caused by data entry mistakes like comma errors. The distinction between parametric and non-parametric statistics is taught based on the underlying data distribution, ensuring robust conclusions are drawn from box plots and histograms. Ultimately, the course synthesizes these diverse topics—from the molecular mechanics of DNA and RNA to the statistical frameworks of GWAS—to provide a comprehensive understanding of how bioinformatics integrates computational power with biological insight to solve real-world problems in health, agriculture, and evolutionary biology.
Read the full video transcript
welcome everyone um last lecture bioinformatics um the summary so if you've not watched any of the previous videos this is going to be the video for you because i'm going to summarize everything in like two hours um so it's gonna be good it's gonna be good so last stream last stream i'm it's like puppy eyes it's such a shame such a shame it's not gonna be the last stream because the next um lecture series is going to be online as well um i hope so i hope so it might be that actually the university will force me to do it in person um but we'll have to see we'll have to see but i'm i'm very sad about it like i like the like weekly being here and and talking to myself and reading chat and answering questions and these kinds of things but um it's really a shame that um we're done with this series so but you guys had 14 good lectures um and uh i hope everyone does really well on the exam so that uh that we don't have to have a makeup exam or something like that anyway this is going to be the overview right so the overview is just going to be me talking through all of the lectures and then all the way at the end i have four example exam questions they're actually not example exam questions they're actually real exam questions from previous years so just so that you guys get a feeling of what kind of questions i i ask and how i ask them and i will be answering them so then you can see what i think is important but with that out of the way let's start with lecture one the introduction um so in the introduction we talked a lot about what bioinformatics is right like bioinformatics is a discipline that uses tools from computer science to answer biological questions and then i also gave you guys a whole bunch of this kind of definitions but the thing is that if i would ask a question like what is bioinformatics right then i want you guys to at least mention computer science and biological questions right because those are the two core elements in the in the definition so you can write it down any way that you want as long as it mentions computer science or information technology or some kind of analogy for that and of course the biological questions need to be in there as well so that's kind of the way that i kind of do the slides right many of the slides they have this kind of blue highlight in definitions and that that's the important part so for lecture one know what an algorithm is know what it know what data is know what knowledge is know the difference between data and knowledge and also know the difference between like in vivo in vitro and in silico because those things are are pretty important for bioinformatics we also talked about dna sequence right that sequence is the fundamental data type in bioinformatics that that's the thing that started it all right so we started doing like protein sequencing and dna and rna sequencing at a certain point and that's where the whole field of bioinformatics comes from um yeah so dna rna and protein sequences are more or less the the reasons why the field of bioinformatics exists because people needed to store it in databases and then you need to analyze it and know that sequence is also the entry point for many insidico studies right if we think about the current coronavirus situation then and it all started with people sequencing the virus figuring out that it's a cyber cough and then making like phylogenetic trees and tracking how the virus mutates across the world and spreads across the world and had know what whole genome shotgun sequencing is and also know that sequence alignment is one of the most fundamental algorithms in bioinformatics so had the alignment of two sequences against each other and we spend a whole lecture on that so we also went quite quickly through the microarray workflow in lecture one um so and know at which point you don't have to be able to reproduce the whole thing right like i don't expect for you to like know point by point by point what the microarray workflow is but i want to be able to ask a question like um in which parts of microarray analysis is a bioinformatician involved right and then you could say well biophysician is involved in creating the arrays but also in data storage data normalization and and generally i will ask something like give me uh four steps or four things or three things right so all right then lecture number two phenotypes so phenotypes we talked about qualitative properties and quantitative properties so qualitative is something that is um something like it tastes good it smells bad i don't like it or i do like it right so it's something that you it that is really hard to kind of put a number on so and not measured with numerical results and then we have quantitative properties um and and quantitative properties that are properties that exist with a magnitude or multitude which means that they can be measured using s i units and we talked also about mendelian traits and complex phenotypes so mendelian traits are traits which are caused by a single gene which causes the difference in a phenotype while complex phenotypes are phenotypes which are controlled by many genes so one example of a mendelian trait is earwax there's dry earwax and there's wet earwax and there's a single gene in the genome that controls if you have dry or wet earwax right so it's a single mutation in in a single gene complex phenotypes are things like human height intelligence and all of these things right that's determined by many many different genes we also talked about like this mixing flowers thing um so if a phenotype is additive right so if the genetics underneath the phenotype is additive then you get mixing so that means that when we have a red a red flower and a white flower right so these are the gametes then you get this following mendelian inheritance diagram while if we have dominance right so one of the phenotypes dominate then we get a different proportion and this is because one of the two and they can't mix together so if you have a red allele you will always be red also be able to read these kind of diagrams so um yeah i might ask a question about a diagram so i will show you a diagram and then ask you is this a additive or a dominant phenotype furthermore we talked about the concept of linkage right so because genes are located on a chromosome hey the closer they are together the more often they are inherited together so if they are very far apart then there's a high chance that these two phenotypes there will be a recombination in between so homologous recombination when the gametes are produced separating the two phenotypes from each other and so linkage is a very difficult concept and i just want you guys to kind of be able to tell me in your own words what linkage is and we also talked about two-point and three-point crosses which are very much the same um but you these are used to determine if genes are linked or if they're independent right so if they are on the same chromosome and how close they are on a chromosome and then we have independent which means that gene one is on chromosome one and gene two is on a different chromosome for example chromosome 11. and the advantage of using a three-point cross compared to a two-point cross is that in a three-point cross you can also infer the order of the chromosome right so it allows you to kind of build a genetic map where you can say well if we start at the beginning of chromosome one then we first see the phenotype for for example broken wings and then we see the the phenotype for eyes and then we see a phenotype for antenna right so we can determine the order of the genes on the genome and that is only possible when we use a three-point cross because we can kind of figure out if a is closer to b than it is to c and these kinds of things good we talked a little bit about phenotypes in lecture two as well so we talked about visual and analysis like box plots and histograms so be able to kind of tell me things about box plots or histograms right so a box plot generally shows the median value and then it shows the uh the quantiles right so 50 of the data and then 95 of the data generally in the vexes um and then we have like things like uh histograms that we talked about but that probably really would be a question about that we talked about multiple testing so definitely know the difference between a type one and a type two error and we also talked about descriptive statistics right what is an outlier and how to deal with outliers so you can you can winsorize them away head generally outliers are values which are very very far apart from the distribution and they can be caused by things like comma failures where you write down the comma wrong so instead of writing 3.0 you write 30.0 right so these kinds of things happen we also talked about things like exploratory data analysis and to decide which model to use on the data a little bit so if you have a really nice normal distribution and of course you want to go with parametric statistics but if you have like a lot of outliers in your data then it's probably better to switch to non-parametric statistics and a little bit about hypothesis testing so good and then in lecture three we talked about dna right so we talked that dna is used for diagnostics a lot used a lot in biotechnology and forensic biology and in virology right so dna is used to catch criminals uh find things like do you have the bracha gene and have do you have a high chance of developing breast cancer and these kinds of things but also dna and dna research is used a lot in biotechnology right if we want to make like algae produce fuel then also we we look at the dna of these algae and try to optimize them for producing biofuels virology i think that speaks for itself so we also talked about the old more or less classical ways of sequencing dna data right so and we talked about maximum gilbert sequencing pyro sequencing summer sequencing and next generation sequencing and what i want from you guys is that you are able to read these kinds of plots right so here you see a plot which is a maxim gilbert uh no this is a pyro sequencing yeah i think this is pyrosequencing and so you add the nucleotides in order and then have when the nucleotide gets incorporated you see a little flash of light and the height of the of the flash determines how many base pairs there were but hey if if i would show you a figure like this and i would say to you guys like this is the order in which the nucleotides are added then and i want you guys to be able to say okay so this is then the resulting sequence the same thing for this um so maxum giber sequencing where you where you use like where you use kind of cutting enzymes which cleave at different points so you have four different cutting enzyme one which cuts at a plus g one which cuts at a g one which cuts at a g and one which cuts you the c plus t which causes to which causes the dna to fragment right and then these fragments are brought up on the gel and then one of the things that i saw a lot in recent years when we did the exam is that people actually read it the wrong way around so the sequence here you read from from the bottom to the top right because here we see that it's c t a c g t a and here you see c t a so hey you don't read it and that's that's what often goes wrong so people are able to kind of figure out which base pair there was at each of the different positions but of course the first position is the lowest one because the cutting enzyme cuts at a certain point right so the smallest fragment is the is the is the first base pair so remember that when you see a sequencing gel with maximum gearbear sequencing you have to read it from the bottom to the top and that that goes wrong a lot so that's a tip for you guys we talked also about the workflow for next generation sequencing and that you do sample preparation then you do dna sequencing generally as a bioinformatician you're not involved in that sample preparation is done by a postdoc or by a phd or by a master's student in the lab and the dna sequencing is generally done by an external company because like at academia we almost never do our own sequencing but of course have what you get from the company is these fast q files and then we need to do all kinds of steps before we can end up with a list of our single nucleotide polymorphisms so here we need to trim the reads which means get rid of the ends because this the quality of sequencing drops so the more base pairs i sequence the lower my confidence in that the base pair is actually correctly sequenced so had a certain point we decided or in the workflow you have re-trimming where you say well i have my read my read is 150 base pairs long but i see that the quality after like 110 drops off so then i'm just going to cut the read there and i'm going to ignore the last 40 base pairs so the re-trimming again yields a fast q file same format as we had and then we do alignment and alignment is just taking the read scanning across the genome see where it fits and then here you get a bum file which is this kind of file format used for um for next generation sequencing data which is similar to the sum but then binary but after alignment we have to handle duplicates because in the sequencing process we have generally a pcr step when we do our sample preparation but also we have optical duplicates which are caused by how the machine works so we have to remove duplicates which means that if we have a read which is starting at a certain position ending at a certain position but we see the same read over and over and over and over again then we we just ignore all the duplicates and we just say no we had one read at this position instead of having like a hundred or a thousand right so these optical duplicates are very common so you have to remove those then the next step is indel realignment so indel realignment means that you use known variation in the genome right we've already sequenced hundreds and hundreds and hundreds of humans probably more in the order of hundreds of thousands of humans by now so if we do an alignment of a read towards the human reference genome then of course it might be that inside of where the read fits that there is a variant right so a variant means that there's a single base pair which is different in some individuals but that means that we don't want to penalize the alignment for having this variant so that's what indel recalibration does head looks at little insertions and known deletions and then says well i'm not going to penalize the read for this because this is a known variant in the human genome and of course we don't have it just for humans but also for mice and rats and other model organisms and then we have the base recalibration step so the base recalibration step is very similar to the indel recalibration step or the indel realignment step but the base recalibration step is just looking at single nucleotide polymorphism right so it the indels are for kind of short deletions and insertions and the base recalibration is the same thing but now for for single base pair variants and then in the end generally what we do with dna data is we don't look at the whole genome that we have but imagine that we have a human then we kind of want to summarize where the human is different from the reference sequence so we then do single nucleotide polymorphism calling or snip and indel calling to find the regions or the yeah the positions in the genome where our sample is different from the reference so also know that there are drawbacks about doing next generation sequencing right you need a lot of computing time it's getting better and better because tools become better and better of course but in the end there's a lot of computational time involved in doing the analysis of dna sec data you need a lot of hard drive storage you need a lot of random access memory and you need management of files so hey you need to keep track of all of these different files that are being produced because of course in the whole pipeline we start off with one file we get from the company which turns into two files three files four five four six seven so it's like seven eight files that you have in the end and you need to manage those and those need to be stored and you have to have backups and these kinds of things also know that there's a difference in the definition of what a gene is right in the previous lecture so in the lecture when we talk about genes or as in phenotypes right so units of inheritance and but in molecular biology and in sequencing a gene is not a unit of inheritance a gene is a union of genomic sequences encoding a coherent set of potentially overlapping functional products right because a gene nowadays in biochemistry or in in in molecular biology we see a gene as having introns and exons and a single gene can produce different proteins or different variants of the same protein and so it's a much more complex definition in molecular biology than it is when we talk about genetics in genetics a gene is very basically a unit of inheritance so generally it comes in two forms you have a red gene and a white gene and hey you get one of the gametes from your father one of them from your mother and these mix but of course in molecular biology stuff becomes much more complicated because you have a gene which encodes color and this color gene can have like 10 different variants right some of them are white some of them are red some of them are purple some of them are blue and all of these things can mix and match in an additive or in a dominant way with each other so in molecular biology a gene is is a very um it's a fixed definition and and it it's a it's a very it's a very good definition but it's a different definition than than what we use in genetics so be aware of that we also talked about transposable elements there will definitely be a question about transposable elements remember that they are that they were first described or they were discovered by barbara mclintock one of my favorite molecular biologists ever so there might be a question about that but when we talk about transposable elements so also called jumping genes they come into two different classes so you have retrotransposons and dna transposons so the retrotransposons they have an intermediate rna form so they are more or less in the dna they get more or less transcribed into rna the rna gets then built in into the dna as well and the class ii transposons are dna transposon so they don't have this intermediate form and then every one of these classes is subdivided in two so you have anonym autonomous retrotransposons and you have autonomous dna transposons and autonomous means that they don't need anything to move so everything that they need to move from one position in the genome to another or to copy themselves from one position to another they carry with them non-autonomous means that it needs something from the host cell to move from one position to the other one right so it means that not all of the proteins that it needs to jump around are encoded on the transposable element themselves we also talked a lot about different regulatory elements so just read through them and know that there are different types of regulatory elements like insulators enhancers you have tata boxes and and you have like metal sensing elements in the dna um but i'm not going to ask in too much detail about that i think that the transposable elements i like that much more so there's more likely to be a question about transposable elements than there is about regulatory elements and also know the difference between a mitochondria and a chloroplast know their function hence the mitochondria are the powerhouse of the cell which means that they produce atp chloroplasts are the same thing they are also the powerhouse of the cell but then in plants right so they they do photosynthesis and produce atp for the plant that way lecture four rna so i think this is generally the most boring lecture for everyone also for me because there are so many different like types of rna right so you have messenger rna which comes in pre-messenger rna called hn rna and then you have the mature messenger rna called mrna we have transfer rnas which transfer amino acids so they form the the link between the the messenger rna sequence and the protein sequence by by having well on one side they have the codon and on the other side are on one side they have the anticodon right which matches the messenger rna and then you have the amino acid which is attached to it in a cloverleaf system you have ribosomal rna which is rna inside of the ribosomes which helps the ribosome be able to produce proteins we have small nuclear rnas which are in the nucleosomes in the in the nucleus which do things like splicing we have catalytic rnas like ribozymes which have a function themselves so it's proteins generally have a biological function but catalytic rna like ribozymes they also have a catalytic function so they are involved in biological processes we have micrornas which are there to do regulation of gene expression so hentai generally are binding messenger rna which then gets degraded because the cell doesn't like double-stranded rna we have small interfering rnas which is kind of a micro rna which is brought into the cell by humans or by micro injection and we have non-coding rna which is actually rna which does not code for a protein but we don't know exactly what it does right so generally it's like the it's like the micro rnas but then much longer so long non-coding rnas or ncrna so there are a lot of different types of rna a lot of different types of definition i don't really like to ask very specific questions about it but i do think that it's important that you know that like rna is divided into all kinds of different subgroups again the workflow for rna sequencing is the same as for dna sequencing the only big difference is is that you acquire your samples you extract the rna instead of extracting the dna and there is this additional step where you do rna to dna reverse transcription and of course in the end because we do rna sec we're not in we're generally not interested in the snips so the variations in the genome we are generally want to do the extraction of the expression levels at the end right so instead of saying that well at this position my sample is different from the reference genome in this case what you are going to do is say well i look at my gene of interest and i count the number of reads that are there and then i'm going to take the number of reads in sample 1 and compare them to the number of reads in sample 2 to see if there's a difference in expression level so head the goal of rna sec is different from the goal of dna sequencing in that you want to get the expression of the genome so the expression of the different genes in the genome while generally in dna sequencing you want to look for variations in the genome we also looked at tools to predict secondary structure of rna right so generally you take the sequence you annotate groups of secondary structures and then this is all based on the lowest free energy structure right so it tries to fold the rna in such a way that there's the least stress on the molecule so here we have things like rna fold which i think we had an example of but there's also context fault and rna shapes and there's a lot of different tools in the rna lecture we also said that if you look at these short rnas or if you look at these long non-coding rnas right they generally have like this modular structure which means that there's hey if you have a long rna molecule then part of it can for example bind rna or dna part of it can bind proteins but also parts of rna can be conformational switches and these things they are buildup modular right so you can have a long non-coding rna which has two conformational switches and a protein binding domain or you have a long non-coding rna which has a dna binding domain a conformational switch and then an rna binding domain so and based on on which which kind of structures we find in the rna we can kind of figure out what the function of this rna is right if an rna has a protein binding domain right a piece of the rna is predicted to bind the protein then of course we can kind of infer that this rna has something to do with proteins so but it's a modular structure so these long non-coding rnas they are modular so they're build up different modules which are more or less mixed and matched together all right so in lecture five we talked about proteins so here we have some nomenclature right which i want you guys to know so an amino acid is a single building block we have a polypeptide which is a chain of several amino acids and then we talk about an apple protein which is one or more polypeptides but not having the cofactor so for example the zinc molecule that is needed to bind the thing that it needs to bind or the iron molecule to bind oxygen when we think about hemoglobin and then when we talk about proteins right then we talk about apple proteins with cofactors right so hemoglobin is a protein and then when we say the hemoglobin protein we mean the four chains or the eight chains of hemoglobin i think it has four so it has four chains so there's four apple proteins so four of these polypeptides and then within these polypeptides you have iron molecules which bind oxygen right and then we talk about a protein so we also talked about hirality right so the fact that if you have an amino acid right then almost all amino acids have are chiral right because this molecule in 3d right cannot be put on top of this right it's the mirror image right that is what chirality means that you you have a molecule and then you have the mirror image of the molecule and these two although they have the same structural formula they do not have the same 3d structure and because of that you can have one of them being very toxic and the other one being very beneficial right so we also talked about that in nature most of the amino acids are found in the left form so the l form and the d form is generally not seen or it's generally not produced but this the chirality itself in proteins becomes a big issue when you do like um chemical synthesis of um of of medication right because when you do chemical synthesis uh the chirality is um egal right because we don't care about the chira or the the the process the chemical process that we use um uses um like a plus b is c right but when c is produced it's produced in both forms so when you talk about amino acids remember that they are chiral also remember that there is one amino acid which is not chiral and that is glycine because glycine has an h as the r group right the r group is the kind of side chain which determines which amino acid we're looking at and of course when r is an h then we are able to turn the molecule in such a way that we we end up with the mirror image so glycine the smallest one so when the the side chain is just a single hydrogen molecule then it is not chiral um you can draw it and then try to uh try to do it there's also these boxes actually um so you have these snappy atoms snap ems or something like that they're called and there you can just build these amino acids right so you have c molecules and you can stick in the things um so if you if you are interested in chirality and stuff then then pick up one of these boxes of snap snap or snap atoms i don't know exactly what they're called but then you can you can build these atoms yourself which is really fun so when we talked about proteins we talked about the fact that you have the primary sequence so the primary structure is just the amino acids in a row right so you have glycine violin volume lysine isolyzine so when we talk about the secondary structure the secondary structure and the primary structure of course and this is what i pointed out in the lecture um is base the primary structure of proteins is based on atomic bonds and because of the fact that some amino acids actually are able to form sulfur bonds you can have primary structures which are not just a single line of more or less letters right you i showed you guys i think i showed you in the lecture two or three more or less complex primary structures where you have two polypeptide chains which are connected together by a sulfur sulfur bond um because of the and and that is that is the difficulty in primary structures for proteins is that unlike dna and rna which is just a single more or less straight line of letters um in proteins the primary structure has already other interactions and and so primary structure is based on atomic bonding secondary structure is based on um hydrogen uh bridging right um so and then we have the tertiary and the quaternary structure so the tertiary structure is based on more or less all forces working on it and quaternary just means that we take the whole protein so the different polypeptides note that there are different computational tools to predict protein structure so there's up initial prediction where you just take the primary sequence and then try to predict secondary tertiary and quaternary structure but we also have dedicated tools for secondary structure prediction because that is more or less something that we can do very well but from the primary structure determining the tertiary structure is really hard there's really good tools out there which can actually predict if there will be an alpha helix and if this alpha helix will go through a membrane because these things are very are very common right so we know exactly how transmembrane alpha helices look like there's thread and fold recognition and homology modeling so hen know that there are five different more or less schools of thought about how to predict secondary tertiary and quaternary structure of proteins from primary structure we also talked about the new alpha fold from google which is kind of using machine learning to do it but again machine learning is just the field of homology modeling right because machines they look at all kinds of examples and then they learn how a protein folds based on the examples but that of course is kind of a type of homology because learning from an example means that you use homology we also talked about how you can separate proteins right so head there's two dgl electrophoresis which allows us to separate protein mixtures and we separate using two different methods so the the standard method is used for the y-axis or the the the y component of the gel right so that's the same for rna gels dna gls and and protein gels and so here we separate based on size using an electric charge and then in the other axis on the x-axis we separate using a ph gradient so we start off with a very low ph low ph of like two and then we end up here with a high ph of like 14 right so water is like seven so in the middle so every protein comes with a charge and that is because they have side chains and the side chains they give this protein an intrinsic charge which means that a protein which has a positive charge feels more at home in a negative environment right and a negative environment means that you have an abundance of hydrogen so that means that you are then in a positive ph but i could be wrong right but had this this second axis is based on the isoelectric point and the isoelectric point from a protein or a protein isoelectric point is made because of the fact that the protein has side chains we also talked about orthologs paralogs in parallax out parallax center locks and i want you guys to kind of know what it is um and i i hope i explained it well but it was at the end of the lecture after like a two and a half hour stream so if i didn't explain it properly um in the in the lectures then do look it up online because it is important there will definitely be a question about what is the difference between an ortholog and a paralog right so and this has to do with the gene duplication events and speciation events and so when a species splits into two species or when a gene duplicates itself across the genome and i hope i explained it well during the lecture but if i didn't then please look it up because like i can't explain everything perfectly because if i've been streaming for two and a half hours then sometimes the quality of my thinking goes down uh and besides that we have of course xenolog so xenologs are more or less pieces of dna or proteins which are transferred from one species to another right so it's a horizontal gene transfer mechanism and we we talked about like four of them and one of them of course is just cloning or genetic engineering of bacteria but bacteria also exchange dna with with other bacteria so they make these little tubes and then they just exchange parts of their dna with each other to increase survival for both of them all right lecture six was about metabolites so we talked about endogenous and exo endogenous metabolites and exogenous metabolites heso know the difference between the two we also talked about primary metabolites and secondary metabolites so primary metabolites mean that if you don't have them you more or less die instantly while secondary metabolites are metabolized which you can go without so we also talked a lot about the mass spectrometry workflow so mass spectrometry is four different steps the first step is compound separation which can be done using three different techniques two of them which are chromatography techniques either using a liquid or a gas as the mobile phase and then we also have capillary electrophoresis which again is very similar to how we separate proteins and how we separate dna by their size but here we use electrophoresis using a very narrow capillar and in the capillar we kind of break down or we we we slow down big proteins because they are big and and small proteins go through relatively quickly so after we've done the compound separation in mass spectrometry we go to fragmentation and ionization which means that the the protein that or metabolite that we're looking at gets fragmented into little pieces and then each of these little pieces gets ionized so they get a charge put on them so this can be two positive charge or three or four or one positive charge and of course we do this to be able to have the the the thing fly through the mass spectrometry right because it needs to be charged to be attracted or to be shot out and then we have the separation of the mass over charge right so we can do this using a sector instrument or a time-of-flight instrument and then we have the detection so the detection part is actually just generally a the the charged molecule molecule flying against the metal plate and then this is detected using a computer we also talked about keg so we talked about headed keg as pathway information and that it's based on kind of a input protein output right so you have a metabolite and then a protein working on that metabolite transforming it into another metabolite right so it has these kind of compounds and reactions um hey so they have genes and and proteins in there but the main selling point or the unique selling point about keg is that it allows you to reason what type of metabolites an animal can make and which type of metabolites the animal cannot make we also have reactom which is a different database which is very similar right it also contains pathway information it also has many different organisms this one is open source cac actually has a paid version and also a free version but the difference between keg and react home is that cag is very much based in more or less chemistry right so we have a metabolite and a protein working on a metabolite transforming it into something else while reactome is more holistic in a way right so they have a pathway for rna or for rna transcription or dna duplication right so they're they're they're pathways in reaction are very similar to the pathways in keg but they look at a slightly higher level so it's not metabolite protein metabolite it's it's it's more conceptual we also talked about cytoscape one of these open source tools that allows you to visualize complex networks and integrate with any type of attribute data which of course means that it's used a lot in bioinformatics to show like large gene networks or large protein net networks but it's also used in social network analysis which means things like facebook and you can use cytoscape to visualize your friends and who are their friends and hey you can then use different attributes so you can say well everyone living in germany color them green everyone living in in poland calling them blue and all of these things right so you can overlay all types of different data on top of your network and that is why cytoscape is really useful and we also have it is also used a lot in the semantic web so have when websites are presented to you it's just plain text but you can use html tags to html tags to kind of give meaning to parts of the text right so you can tell for example the search engine saying that denny is a name right and it's a first name while aaron's is a family name and then the search engine starts to understand what's going on and it can build up kind of an internal network saying that okay so denny adams is a person and he works at this department this department has other persons working there and then it can kind of form a more comprehensive image on what is being displayed on a website and this is called the semantic web or web 2.0 i think and nowadays people are talking about web 4.0 i i got lost at web 3.0 like for me it's all html cms and javascript but there's there's a apparently a difference between the world wide web now and like 10 years ago i think the main difference is just that the spying is increased a lot so um then we had lecture number seven the introduction into r and there will be no questions about this on the exam because this is just a lecture for you guys to show you guys that you should if you want to have a career in bioinformatics you should definitely pick up at least one data science programming language so hey it's really good to learn something like r or python which are more or less the two main languages that are being used in bioinformatics so but for you guys there will be no questions about this on the exam so that's good that means that you can just skip the lecture when you are learning for the exam all right and then we went all the way back right because now we discussed all of the different biomolecular levels we started off on the lowest level which is the dna then the rna then the proteins then the metabolites right and then we started talk i started talking or the lecture eight was about phenotypes and how we do qtl mapping right so i talked to you about the quantitative traits and the i uh the ec also a little bit of repeat of the first lecture qualitative traits which are more or less measured subjectively i showed you this picture where we say that quantitative trades are a subset of all trades out there so all trades out there are qualitative and quantitative together but of course like quality quantitative trades are a subset of the qualitative trades and this subset is growing right the more machines we build the better we are in kind of um expressing qualitative things into quantitative units right so an example of this would be uh the the taste or the quality of wine that used to be a very qualitative trade right you would have a panel of wine tasters everyone would taste the glass and then they would score the wine saying this is a good or a bad wine but nowadays you just have a robot that does that right so a couple of drops of the wine get put in the robot and the robot analyzes the composition of the wine and then just gives it its score so quantitative is growing while qualitative is more or less shrinking and again i talk to you guys about mendelian and complex phenotypes so when we talk about phenotypes in qtl mapping i taught you guys about the crossover events right so that we have meiosis 1 meiosis ii and that this whole thing works or this whole thing is um and that we can do things like associate a region of the genome with a certain phenotype or find a region of the genome where a phenotype is more or less controlled from and that is only possible because we have this chromosomal crossovers right so that that in meiosis one has so what we get is we get duplication of the of the of the genomes that we have and then we have the homologous chromosomes which are more or less bound together and then we have this this crossover where parts of one chromosome are exchanged to the other one right and then of course we have the meiosis ii where now we go from having two copies of each chromosome to having only a single copy of the chromosome and then these are called generally gametes so and i also had two links there in the in the lecture which are two more or less little movies where it is explained in much more detail and graphically right because a movie can show you guys how this happens so that they align together and that they then get swapped around and that's very difficult to catch in a in a picture so when we talked about qtl mapping i told you guys that you can only do qtl so quantitative trade locus analysis when you use an experimental cross right so you start off with for example two inbred founders who get crossed together then we get a generation which is called the f1 generation and in the f1 generation everyone has one chromosome from the father one chromosome from the mother right and there's no no crossover here or no recombination because of the fact that the parental line had two exactly identical chromosomes so for the father it had two exactly identical chromosomes and for the mother the same thing so had of course crossover occurred but this crossover had no effect i think i even made a little drawing during the lecture showing how this works but i want you guys to know the advantages and disadvantages of the different types of crosses that we discussed and so for example what a recombinant inbred line is and so where you do this cross between the two founders and then create these funnels in which you within the funnel you start brother sister mating so to make sure that you get immortal animals um which you can use forever and ever um well they're not immortal but they're like clones right so a single real line so recombinant in red line so one of these lines you can mate a male with a female and then the children of these will be exactly identical but there are of course problems there because if you have a recombinant in red line then there's only two states so either being a a or bb so there are no heterozygote animals within the population and because there are no heterozygote animals in a recombinant inbred line you cannot estimate things like additive and dominance you can only see that there's a difference between the two homozygote groups but you get no information about the heterozygotes in the middle the back cross is more or less the same thing but the back cross is really quick to make because you cross two inbred animals you get an f1 generation and then this f1 generation is crossed back to one of the two parentals and then have we have the we have the advantage that it's really really quick to do because you only need two generations but the problem with the back crosses is that when you do the association you get large parts of the genome which are associated and you have this imbalance between having only 25 percent aaa and 75 bb and again because individuals are only a a or a b you get no information about dominance and additivity f2 cross more or less solves these things so it has the disadvantage of still having like large regions but it allows you to investigate additive and dominant effect so qtl mapping gwas we talked about the difference we talked about how they are very similar right so there's there's they're both methods to find regions of the genome which control genes are which control phenotypes right and the the differences is that in qtl mapping you are able to map between the markers because of the fact that you have a structured population while in a g wash you just have an outbred population generally of humans and there of course you cannot know what is between the markers but in a in an f2 for example you can map in between the markers and of course there's another big difference and that's the way how these results are displayed so in a qtl you have a smooth line plot across the chromosome and in a genome-wide association you generally have the results presented to you as a manhattan plot so had during the lecture we saw examples of that i talked also about effect size versus likelihood so that the effect size is the the difference between the aaa and the bb group and that the likelihood is the statistical test when you compare individuals having a a versus individuals having bb and of course here we have to also think about multiple testing but that came back i think in another lecture as well so multiple testing is of course the issues that when you do a lot of statistical tests you have to kind of compensate for the fact that you did a lot of tests good so i've been talking for around an hour so we'll do a quick break and then we will do the remaining five lectures and then go to the um four example questions so that you guys have an idea of what i ask and when um so let me set up not the audio but the music um so yeah i'll be back in like 10 minutes and then we'll just continue with uh discussing the different lectures so five lectures left and then then we're done so it's going to be a very short lecture so i will see you guys in around 10 minutes if you're watching this on youtube then um probably see you tomorrow so bye bye for now