Submind YouTube summaries
Thumbnail for Summary and Example Exam Questions (Bioinformatics S15E2)

Summary and Example Exam Questions (Bioinformatics S15E2)

Watch on YouTube

Video summary

The video provides a comprehensive summary of a bioinformatics course, reviewing key concepts from lectures on PCR primer design, database management, sequence analysis, gene expression, and data standards. The instructor begins by recapping the mechanics of Polymerase Chain Reaction (PCR), emphasizing the critical requirements for designing effective primers, such as appropriate length, GC content, and melting temperatures to ensure specific binding without non-specific amplification. Special attention is given to advanced techniques like multiplex PCR, the use of universal or guest primers for organisms with unknown genomes, and the importance of avoiding repeated sequences in primer design to maintain uniqueness. The discussion extends to the thermodynamic principles governing DNA hybridization, explaining how annealing temperatures must be carefully balanced with elongation steps to ensure stability throughout the cycling process. Following the review of molecular techniques, the summary shifts to the organization and querying of biological data using databases. The instructor explains fundamental database concepts such as normalization, indexing, and sharding, highlighting their advantages over simple file storage for managing large-scale genomic datasets. Specific bioinformatics resources like ENSEMBL, PubMed, PDB, and dbSNP are introduced, along with tools like BioMart that allow researchers to programmatically retrieve data without manual downloading. The lecture also covers the mathematical foundations of sequence analysis, including global versus local alignment strategies, scoring functions for matches and mismatches, and the distinction between linear and affine gap penalties which better model biological insertion-deletion events. Furthermore, the importance of substitution matrices like BLOSUM and PAM in evaluating protein similarity based on chemical properties is discussed. The final portion of the summary addresses gene expression analysis through microarrays and RNA-seq data standards. Key topics include the necessity of normalization to correct for dye variations and experimental artifacts, statistical methods like t-tests and ANOVA for comparing groups while controlling for covariates, and multiple testing corrections such as Bonferroni and False Discovery Rate adjustments to manage Type I and Type II errors. The instructor also reviews file formats essential for bioinformatics workflows, including FASTQ for raw sequencing data, GFF for genomic features, VCF for variations, and PLINK/BEMAP for association studies. Additionally, the video covers clustering algorithms like single, complete, and average linkage, as well as distance metrics such as Manhattan, Euclidean, and Minkowski distances used to analyze expression profiles. To conclude, the instructor presents example exam questions covering fundamental biological knowledge, such as the difference between pre-mRNA and mature mRNA regarding intron splicing, the function of tRNA in linking RNA sequences to amino acids, and the four main steps of a mass spectrometry workflow. Practical advice is offered on protein purification techniques and software development practices, including unit testing, regression testing, and test-driven development to ensure code reliability. The session ends with administrative details regarding the upcoming written exam format, registration logistics for different student groups, and an announcement that future data analysis streams will resume in April, encouraging students to focus their study efforts on the core theoretical concepts rather than introductory programming modules.
Read the full video transcript
back everyone if you're watching this on youtube thank you for being here if you're on twitch then also thank you for still being here um so let's just continue so lecture nine was about primer design so the first part was about polymerase chain reaction right so [Music] what do we need for pcr so we need water nucleotides primers template dna and a master student to do it for us but we also talked about what is a good primer and when is a primary primer right because primers need to be a certain length and they need to have a certain binding capacity by having like an ac gt kind of uh so an a atgc composition and we also talked about advanced primers so that you can do multiplex pcr where you have or where you're not amplifying using a single pair of primer but you're using multiple pairs of primers we talked about universal primers and semi-universal primers if you want to amplify a piece of a virus but not just from one strain but multiple strains so then you can use things like universal primers i think the example that we did there was happy fe something like that but then here we also have guest mirrors so if you don't have a dna sequence available then based on a known protein sequence you can still design primers for your animal which doesn't have a genome sequence available and you do this by back translation right so you look at the amino acid sequence and then you code the amino acids back to their dna equivalent which of course is not perfect so gusmar primers are generally longer than standard primers to make them still bind or have to still give them this the binding properties that you need so when we talk about pcr pcr comes in three steps so know that the first step is denaturation where we heat up our sample to go from double-stranded dna into single-stranded dna and then the next step is lowering the temperature a little bit so that the primers can bind so ahead this is generally like 90 degrees celsius annealing happens at around 60 something degrees i think the the the temperatures were mentioned in the lecture um and then had the primers bind to the dna and then in the next step we put the temperature up to like 70 degrees celsius and then the the the dna polymerase starts amplifying the the dna and dna polymerase only starts amplifying dna when when we have these primers there so when when there's double-stranded dna so we talked about pcr right so when we have our gene of interest that we are trying to pcr out then of course in the first cycle we get like two copies four copies eight copies and sixteen copies um let me just mute myself i don't know what's going on with my voice but um so hemp in in in pcr we do an exponential amplification which means that the number of cycles that we use we can kind of estimate how much uh product we're going to get which is 2 to the power of the number of cycles that we do we also showed the first few cycles in detail right because this is just a kind of stylized picture and this is what we would like to happen but of course this is not how it exactly happens but if you're interested in that then please watch back lecture number nine um so how we we also talked about the length of the primer and the uniqueness and the fact that when the primer gets longer you need to have a higher melting and annealing temperature right but there's this trade-off because the longer your primer the the higher the chance that it's unique but also the longer the primer the higher the annealing temperature and then you you end up in this situation where there has to be enough temperature difference between the annealing step and the elongation step and have we talked about things like melting temperature so the melting temperature is the temperature at which half of the dna is single stranded and half of the dna is still double stranded so the melting temperature for genomic dna generally tends to be at around 93 degrees celsius but of course for for primers this is much lower because primers are a lot shorter and we also talked about annealing temperatures so the annealing temperatures the temperature at which the primer starts binding to the genomic dna have we talked about multiplex pcr semi-universal primers guest mers and many more and i also showed you in which fields primer design skills are required right so if you are ever going to do real-time pcr or studying population polymorphisms using microsatellites or aflp markers and then you you use primers and we also talked about internal probe design but then the most important things to to remember about this whole lectures is that if you design primers you have to achieve the appropriate hybridization specificity so it has to be unique right you only want to amplify one part of the genome and you have to be able to do this stably which means that you have to have enough difference in the different temperatures for the whole process to be able to kind of cycle through these three different temperature levels right because if the primer is too long and your annealing temperature is around 70 degrees celsius right then the annealing temperature is the same as the temperature used by the polymerase and then stuff starts going wrong because then then the stability doesn't hold and stability also has to do of course with the number of ats versus gcs because you want to have your primer to bind very tightly to the dna and of course this has to do with the template dna and the ac gt content of of the of the template dna as well to remember right if you design primers never design primers on repeated sequences um because if you design primers on repeated sequences then of course they're not going to be unique so primers cannot contain repeats themselves either and so we want to get rid of them and for that we can use a tool like repeat masker right so you can use repeat masker standard or you can use it directly in ensemble so when you export your sequence from an example you have the option to mask your sequence which means that it kind of blocks out areas which are repeated across the genome and it blocks out areas which are of very low low variability right if you have atc atc atc atc then it will block out this region so that you don't design primers by accident to these kinds of sequences so lecture 10 we talked about databases so we talked about some terminology about databases so what is things like sql and these kinds of things i try to explain to you guys why we need databases because they have advantages over just storing data on on your hard drive and they allow things like sharding which makes them much faster and you can do indexing and so now sharding means that you put it on different sites so you have one database which is not just local on one position on the earth but which also has for example the same database but then in japan so that people in japan can also use it and then you have indexing so indexing means that you you look through a column see what's there and already start anticipating on future queries and by building an index so that you can quickly find back stuff in the database we talked about normalization of data within a database and the organization of a database so that you can search generally in different ways and had that you have and we we had an overview of all kinds of important databases in bioinformatics and biology so you need to know that what can we use ensemble for right what what is in pubmed but also that there are databases like the pdb which focus entirely on proteins that you have dbsnp which only focuses on storing single nucleotide polymorphisms in humans and during this lecture we also talked about biomart and i did a biomart example using our african goat data i think and hep biomart is a connector for r and for other programming languages so that you can automatically query many of these important databases in bioinformatics so that you don't have to go to the database to the website and start clicking on it and downloading stuff no biomart allows you to write code which will retrieve data for you which of course makes your research much more reproducible so a little bit more about the normalization right so normalization of data means that instead of storing the full student name in a single column and you can break it down to a more granular level saying that no every name has a first name middle name last name right and then that that's easier because now we can see that all of the students have the same middle name which of course is not obvious from the data that we used to have in in a single column i told you about over decomposition so when you start chopping up things that actually should not be chopped up right when you have a phone number then store the phone number in a single column don't start doing smart things like storing region codes area codes and then phone extensions because these might change over time plus that no one's interested in no one's ever going to do a query give me everyone who has this phone extension right because that's not a very good thing we also talked about things which go wrong in in databases right so for example duplicate data and not so much duplicate data but the worst thing is duplicate data which is named differently right so if you have fifth standard and fifth standard once you write it with the let up with the number five th and the next time you write it with fifth then of course um the database does not know that these two things are the same right so it is a duplicate data but it's even worse because it's duplicate data but it's also inconsistent and this is the thing that databases help you solve by using things like foreign keys right so in this case you would say no the standard is a foreign key which points to another table where all of the different standards are described so we had a bunch of slides about this and and you don't have to know all of the normal forms but know that there are things like over decomposition and that that generally the the database functions best when data is more or less stored in a normalized fashion all right so in lecture 11 we talked about sequence analysis so how we we know that sequences change during the course of evolution and so you have point mutations where a single base pair gets modified and but we also have insertions and deletions and to make it even worse we also have this sign of xenologs and so genes which are transferred from one bacteria to the other but head generally these are considered the three main forms of of sequence variation uh during evolution and then of course we talked about a homology trick right so that since all of all life on the planet more or less comes from one single event where life existed and then everything branched out and we can use homology to kind of infer the function of a protein right if we know that a certain protein is transporting oxygen in humans and then if we find a protein in in mice which has a very similar which is a similar very similar sequence and then we can also assume that this protein and mouse probably also star uh will transport oxygen right so that's endomology trick works because it's a single tree which kind of grows up from from the early beginnings when we talked about sequence analysis we talked a great deal about sequence alignment right so because that's kind of the fundal fundamental algorithm in in bioinformatics and that there are two major variants of alignment one is global alignment and the other one is local alignment so global alignment tries to match the entire string to the other string so it is more likely to insert gaps at the beginning but when we um i think something went wrong with this picture yeah this is this is wrong um please please ignore this slide and and look at the slide in the lecture i think i copied the same one for but her local alignment tries to find the optimal substring while global alignment tries to match the whole thing so yeah yeah so this is this is just wrong i copy-pasted the same thing more or less twice but then look at the original slide in lecture 11 and and know that there's a difference between global alignment and local alignment and local alignment generally is used when you have very short sequences which you want to compare towards the genome while global alignment is used when you have complete viral genomes that you want to align together so we talked about sequence analysis a lot and um when we talked about it we also talked about scoring functions right so that to compare if two sequences are similar we need to kind of have a mathematical definition of similarity right so the the most basic definition that we had or that we could come up with is just the percentage of matches so how many base pairs match and so you get a plus one for each base pair that matches and for every mismatch you you get a minus one penalty um in a way right so percentage of matches is just seven out of twelve base pairs match but you can also add um this this plus one minus one system and this plus one minus one system can then be extended with a linear gap penalty which means that when you open up a gap in one sequence then opening up a gap of two compared to opening a gap of four the gap of 4 is twice as expensive as the gap of 2. but since in biology we know that insertions and deletions are very common we nowadays almost always use an affine gap penalty that means that you get that you have a high penalty threshold for opening a gap but then when you make the gap bigger you don't put that much penalty on there right so opening a gap might be a score of minus one but then going from a gap which is from one y to a gap that is too wide you give a 0.1 penalty right and going from 2 to 3 again you get a 0.1 penalty so and this is this this allows you to do much better alignment when you use this affine gap penalty so know what the difference is between a linear gap penalty and an affine gap penalty and so the the main difference is that in a linear gap penalty you score for every for every x base pairs that you make the gap bigger you get at plus x penalty while in an affine gap penalty you don't have that opening a gap is expensive but extending a gap so making the gap bigger is relatively cheap and i want you guys to be able to calculate the percentage of matches on dna and protein level so when i give you one of these kind of amino acid codon wheels you should be able to read the amino acid codon wheel and if i give you two dna sequences you should be able to say well okay dna wise it matches 7 out of 12 but on protein level we see that there is a 9 out of now a 4 out of 4 right because of course we come in codons so three letters of dna become one amino acid but be able to use these things so i i think we did a small example in the in the lecture as well we also have to take care in dna alignment that we score transitions and transversions differently so because when you when you look at dna then it's we had this figure where we showed that when you go from an a to a t that this is relatively common because the the chemical structure of the a base pair looks very similar to the chemical structure of the t-base pair right so by using electromagnetic radiation or or nuclear radiation um it is very commonly uh it's a very common occurrence for an a to change into a t right but an a almost never changes to a c or not almost never but when this happens it's called a transversion and a transversion is is very uncommon and the same thing happens in proteins because in proteins we have pro amino acids which have very similar side chains so head changing a a lysine by a glycine is a very small change because the because the side chain is the same right and then we have something which is the substitution probability matrix where we talked about the blossom matrix and the pump matrix and these matrices they try to catch this fact that when we have a substitution of one amino acid by a different one so we have kind of a mutation there um it tries to score um some more heavily or it penalizes some more heavily than others right if we have a a positively charged amino acid being changed by a positively charged amino acid then this is a relatively common occurrence but a positive amino acid which gets changed by a negative amino acid that of course has a has a much bigger penalty associated with it know what blast is basic local local alignment and know about cluster w so alignment multiple sequence alignment where we try and align multiple sequences together and cluster w when we talk about cluster w we talk about how to detect conserved residues we can find conserved regions but we can also find patterns in our amino acid structure and when all of the amino acids are for example positively charged head then they don't all have to be the same but then still there might be a pattern saying that for the functioning of this protein it is very important that at this position you have a positively charged residue or positively charged amino acid all right so lecture 12 was all about gene expression analysis again we did more or less the exact same thing as what we showed in lecture one where we looked at microarrays right so creating oligo arrays had nowhere bioinformatics is involved you don't have to know all of the different things right but know that a tiff file is just an image file with these dots on the microarray which can be red green or yellow cell file is this proprietary format from alphymetrix which stores data about microarrays in kind of a compressed way i talked about normalization why do we normalize during microarray analysis and this is uh because the dyes have various varying behavior right the the green dye has a much higher dynamic range than the red dye um but there's also variation during the hybridization like the the surrounding temperature or the surrounding humidity has a big influence on how well your sample hybridizes to a microarray and of course there's variance in the manufacturing if i buy an array now and i buy the same array in like five years then of course the the the quality of the array might be might be different right because the technique gets better so the array that i do today is not directly comparable to the array that i do in like five years so to kind of get rid of these effects right these are all effects that introduce variants into the into the sample and we want to get rid of that and that is why we do use normalization besides normalization of microarrays we also almost always look at the log2 ratio of a microarray and this is because of this varying behavior of the dyes where we want to have kind of a linear scale saying that if i go from zero to one um then this needs to be the same as going from uh zero to minus one right so it's a it's a it's a transformation where we go and then we divide the green dye intensity by the red diet intensity and then we do the log 2 and this is just to prevent the fact from having if you have 1 divided by 2 that is different from 2 divided by 1. and i think the slides actually in the lecture explain it pretty well why you want to use it when we talked about gene expression i showed you guys that you can do it using t-tests but also you can do it using anova test right so a t-test is really nice when you have two groups but when you have for example two different groups two different tissues and you have two different factors for example low concentration of of medicine and high concentration of medicine then you are forced to do an anova test right because an anova test allows you to adjust for covariates um it allows you to put up a model saying that my intensity of the probe is related to the condition in which the probe was measured plus the the thing that i did to the sample plus the type of mouse where the sample was taken from right and you can't do that with a t-test because the t-test only compares two conditions so condition a versus condition b and does not allow you to to control for other factors here again i mentioned multiple testing right so the type one error is calling a gene significantly changed even if it's just by chance so type one errors can be avoided by bonferroni correction and the type two errors is when you say that a gene is not significantly changed or not significantly different between two samples that you have while it actually is and you can only optimize for one of the two so you can say i want to have a minimal amount of type 1 errors but then of course your type 2 errors go up and you can say i want to minimize the type 2 errors but then the type 1 errors go up so that's the kind of trade-off that you have to do and had this one can be avoided by bonferroni correction the type 2 error is by mini hulkberg false discovery rate adjustment we also talked about gene ontology right gene ontology is a common terminology we use to describe the things like the cellular component where the gene is found right so a gene can be active in a nucleus a gene can be active in the cytosol or it can be active outside of the cell so that's exported but also we have a common nomenclature so a common terminology for things like biological processes and molecular function and this allows us to do these over-representation tests right imagine that i do a microarray analysis and i find 50 genes which are different between the two animals that i look at then we can do these tests looking to see if a certain cellular component is over-represented in these 50 genes right if all of these 50 genes are nuclear genes then of course we we hypothesize that there might be something going on in the nucleus but if all of these 50 genes or 40 out of 50 genes are located in the mitochondria head then we might assume that no the mitochondria are the thing where where it goes wrong or where the animal has an issue and the same thing for keg right using keg we can actually it provides these map of different pathways in uh in different species and we can actually overlay our gene expression data onto a cac pathway to see if everything in a pathway is upregulated or if a whole pathway is down regulated based on the tissues that we're looking at again here we talked about similarity right because we have to have a mathematical different definition of what is similar um and of course it's different from when you look at when you compare dna sequences to each other or when you compare protein sequences to each other when you compare expression profiles to each other there are three different distance measurements that you can use right so the manhattan distance is just the absolute difference between the sample one and sample two and then across all of the probes that you measured the euclidean difference is more or less the same but it's not the absolute difference it's the difference to the power of two you add up all of these differences and then you take the square root of the total difference and then the minowski distance is more or less the distance generalization for this where you can choose your own m factor right so an m factor of two means that you have euclidean distance but you can also have an m factor of 3. and why do we sometimes use minowski distance because sometimes we want to put more weight on large differences right because 0.1 to the power of 3 is of course much less than 2 to the power of 3. so by choosing a higher m factor you're focusing more on extreme differences compared to small changes which are globally across but three different distance measurements to express how similar or how different two expression profiles are in animals or in mice or in plants again we want to build a tree right because we want to see which things belong together which things are not belonging together and then we also talked about this clustering um so head there's a difference between single linkage so if you have a group or a group of two profiles and a other group of also two profiles then the single linkage is based upon the two most similar elements within the groups if we look at complete linkage then we look at the two most dissimilar elements and had the distance between the two groups is then based on the two most dissimilar elements and then we have average linkage which is also called op gma and then we look at the distance between two clusters is taken as the average of all distances between pairs of objects x in a and objects y and b and that is the mean distance between the elements in each of the two clusters so average linkage is the best but it is relatively expensive to compute when you have literally hundreds and hundreds and hundreds of elements right if cluster one has a hundred elements cluster two has a hundred elements then you have to compare all of them right so you have to compare one versus a hundred the second one versus a hundred the third one versus a hundred so you you do like a massive amount of comparison so up gma is the most computationally expensive method and that is why sometimes people look at single single linkage and complete linkage because it is relatively cheap because you only have to do one comparison we also talked about where you can get free microarray data so go to gene expression omnibus if you want to get free microarray data to work on and write a scientific publication without spending any money and the same thing you can do at array express the massive advantage of array express is that they have curated reannotated archive data which is a very high quality because someone looked at it and made sure that the sample that was submitted is really the sample that people said that it was and that's not the case for for gene expression omnibus gene expression omnibus anyone can upload data even me so that means that there's no curation going on lecture 13 standards for analysis very interesting lecture i think because have we talked about different biological file formats like the comma separated file but also fasta files so sequencing files um or files holding sequence data then we have fastq which is the standard output for dna sequencers nowaday which contains dna sequence data but also dna quality data we looked at the gff format which is the the format for storing genomic features and we have the vcf format for storing variations relative to a reference genome and we also looked at the bet map format which is a very common format when you do association analysis um so it stores variations on one side and other also phenotypic measurements on the other side so it's kind of the common file format used in association analysis like genome-wide association and and qtl mapping we talked about difference in testing strategies so if you write code as a bioinformatician then use tests to test the code that you've written right so a unit test means that you test the smallest unit so you re you've written a function so you you throw all kinds of different input to the function and then you see if what the function gives you is actually correct based on on what you wanted to do with the function right so regression testing is different because regression testing means that you take the code from someone else and then just throw in data see what comes out and then you start modifying the code but you make sure that every time that you make a modification that you run the test and make sure that for the same input still the same output is produced and then we also had some words about test driven development where you where you develop software using this iterative approach where you say i i want to add a new feature so i write a test that tests the new feature the test initially fails then i start writing code i then run the test again and if the test succeed then i have successfully implemented that feature and i continue with adding a new feature read it so it's it's writing tests and then writing code to pass the test we also discussed all kinds of different types of documentation so we talked about user documentation and documentation which is written beforehand we talked about code documentation and so know that there are different types of documentation for different audiences and so you're not only writing code for yourself but you're also writing documentation that belongs to the code like a tutorial for people that will use your code but you also write things like function descriptions so saying that this function has five parameters and these parameters had the first parameter needs to be an integer between 0 and 100 and so there's different types of documentation for different groups of stakeholders when you are writing software last lecture of last week i think that's also one of the most fun lectures because i always like doing it we talked about citations why do we cite stuff in in science and what's the use of it we talked about things like web of science so if it's not in web of science it's not science we talked about google scholar and research gate and things like h indexes and i indexes we talked about scientific reference management and that you should do some form of scientific reference manager management using a reference manager like endnote or mendeley and then i showed you a difference between distributed and centralized version control right so that's that's it's not directly related to literature management but version control is related to kind of software management right because you you want to be able to go back in time to re-run your analysis and this has to all do with reproducible research right so in theory lecture 14 should have been called reproducible research instead of literature but we talked about citations which are there to make sure that when you claim something that you point to the guys that actually did the research for it have we talked about reference managers which allow you to kind of easily include references when needed and version control is there so that your code is also version so you can go back in time because of course code changes um and sometimes you need to re-run an analysis as if it was 2017. all right so with that first exam question example exam question so if you throw the answers in chat then we can go to the next one and we can all go home early so what is the difference between pre-mrna and mature mrna um i have a sound effect for that like let's do the audio and then just do crickets so i'm just gonna continue this sound effect until anyone answers the question all right question uh answer number one it's not spliced hi by the way yeah hi shannon welcome to the lecture um were you here the whole time or did you just arrive um but indeed indeed pre-mrna still has the introns inside of the messenger rna just forgot to say that's okay that's okay it's good that you answer so um yeah so mature mrna does not have introns pre-mrna still contains the introns so head there's this process called splicing which removes them so that's entirely correct all right next question next question what is the function of trna i'm just going to do crickets again i have more sound effects we can do birds let's do birds for now and then uh so anyone can answer like there should be six people viewing it minus myself and my moderator of course um and then uh we can have an answer to uh what what's the function of trna i actually mentioned it during the lecture i think so it's the link between rna sequence and the amino sequence of proteins yay very good so trna is indeed it it reads the code only in the messenger rna and it it links it to an amino acid so that indeed is the function of trna so it's the uh link between the rna sequence and the very good all right next question what are the four steps in a mass spectrometry workflow experiment let me see um everyone typing typing typing good just that sonic sin doesn't answer all of them like misha come on you know this misha like get your head away from the olympics and answer at least one of the exam example exam questions stop watching the olympics like they will win the gold medals even with you not watching them like no i don't one gold in the pocket okay okay so at least we want a gold medal that's good that's good all right so the four steps are of course compound separation right because we have a mixture so we need to separate the compounds then we need to do fragmentation and ionization then it is separation by mass over charge and then it's detection i think i'm doing it wrong i don't have to know it fortunately you guys do um so that's that's kind of the way that it works right so i already gave the lecture so for the lecture i read up on it but yeah the four steps i think are compound separation fragmentation and ionization separation using mass over charge and then i think detection so that that's kind of the four steps by the way i am relatively strict when it comes to numbers so if i ask you guys for three things and you write down four it is completely wrong if you write down two then you are maximumly allowed to score two out of three points but it's not a guessing game right so if i say um what are four reasons to do x and you write down six then it's completely wrong because i'm not gonna pick which four are correct and which two are wrong or the other way around right even if all six are correct then because i asked for four you did not understand the question extraction separation identification quantification yeah that's what google says but uh that's okay google can say something else i think it's more or less similar right let's just scroll back quickly so metabolites compound separation fragmentation and ionization separation and detection yeah see that's okay all right um last example question i think i had four name two protein purification techniques and describe how they work so that that seems to be a difficult question i don't think that i actually mentioned it here because we did have protein purification in the protein lecture but i didn't i don't i don't think that i mentioned this but uh and of course you don't have to describe now how they work because then you're typing for like 15 minutes and uh but these are the kind of questions that you can expect right very basic questions about the the the lectures that we had um and generally i i like these two-fold questions right so that you name them and that you quickly describe how they work will the exam be oral or written um what do you want it to be because um i'd like to do a written exam but in the pre-function studios are known it says it has to be an oral exam but then the question is because i actually submitted a request to the exam committee to have it changed from an oral exam to a written exam and they actually accepted that but it's still listed in achnas as being an oral exam while i'm actually so it's it's it's probably going to be written because i i think that written just makes more sense i'm fine with either all right good um and i think i looked at acnes and i think you are the only one who registered for the first period for the second period there's actually three people registered so since you are the only one who registered for the exam next week and the other people all registered for the makeup kind of exam or the second exam date um in your case we can we can do whatever you want if you say i just want to have them orally because they were through quicker then that would be fine with me as well but i i will think about it and i will let you know because if two or three other people of course still register you can do a drawing question in an oral exam you actually can because we're doing it via zoom so sonic sin can just sit there make a drawing and show it like that right and the question is still perfectly valid like um draw a platypus right like it doesn't matter if it's written down i can draw a puffer fish that's a that's a that's a that's a that's a challenge we can we can look at that um but anyway yeah i'm i'm still thinking about it a little bit i think legally i could do both um since i did get the uh the the okay from the exam committee to do it written but on the other side you're the only one who registered for the first date so um we might just want to do it orally then because that's going to be a lot quicker and i can be a lot more flexible right if you write it down wrong then i have to kind of say this is wrong but if you if you just say it wrong i can kind of sit there and do like right so we'll have to see i will let you know i will let you know i will discuss also with the other people here because when we do it orally i also need to have a secondary examinator there well i don't need that for a written exam since i just have your written exam but i will i will let you know and i will let you know before this weekend so i will think about it i will discuss with my colleagues here and then i will let you know tomorrow probably because i think people can't register anymore for the first day so i think they can still register for the second date but not for the first date anymore anyway it doesn't matter too much there will be an exam next week and you will do perfectly fine because already here you had three out of four questions so that that's going to be going to be in the direction of like a 1.7 i think but we'll see so all right so um that was it for today and for the whole lecture series so [Music] i discussed with you guys all of the different lectures that we had also which lectures you should not focus on learning so don't don't spend too much time on the r introduction lecture although it's an important lecture because programming is essential there won't be any programming questions on the exam so that's it so we're through i think in total we did 50 hours of streaming 50 hours of lecture um so i want to thank everyone that attended the twitch streams um i want to thank you guys for for being there thank you for attending the course um and like i said um i'll mail everyone who registered for the exam with the details as soon as i get the list still didn't get the list should have gotten the list already from the prefunctural but they're not doing and yeah good luck on the exam and i am i hope that everyone will pass so that we don't have to have a third exam date and yeah i think thank you guys so much for being here and i hope you guys learned a lot of course if you have any questions then feel free to ask and besides that uh could we have that powerpoint that acts as a study guide so you mean this one yeah i will i will upload it directly um i didn't um upload it yet yeah no i will do that directly um good all right then yeah thanks so much for being here um i really enjoyed it um it's nice to stream it like or to be able to do it like this at all um i miss the in-person lectures i like the in-person lectures a lot as well um but i i think this is this is as good as we can do with the current circumstances um so thank you guys for for for being here and spending 50 hours with me uh on on bioinformatics and all of the different topics that we discussed um and uh i will see xanax in at least next week on the exam and the other workers students i will probably see them on the second exam date and that that's it for now so unless anyone wants to get rid of some of their channel points less denny bucks and have me make a drawing then i'm actually going to close the stream early for today and then enjoy learning for the exam and then enjoy the weekend already all right see you next week yes and then uh to all the other people that are still watching bye bye and we will see each other on stream uh let me see let me see because i do have another date for the summer semester so streams will continue or restart um let me see so the data analysis using our course will start on 21 april so the 21st of april there will be a at least probably a stream unless i have to do it in presents or if i can do it in presents um but if it's going to be online then 21st of january february april 21st of april we will start the course so see you then and i hope you enjoyed it i enjoyed it a lot and uh thank you for being here