Video summary
The video provides a comprehensive summary of a bioinformatics course, reviewing key concepts from lectures on PCR primer design, database management, sequence analysis, gene expression, and data standards. The instructor begins by recapping the mechanics of Polymerase Chain Reaction (PCR), emphasizing the critical requirements for designing effective primers, such as appropriate length, GC content, and melting temperatures to ensure specific binding without non-specific amplification. Special attention is given to advanced techniques like multiplex PCR, the use of universal or guest primers for organisms with unknown genomes, and the importance of avoiding repeated sequences in primer design to maintain uniqueness. The discussion extends to the thermodynamic principles governing DNA hybridization, explaining how annealing temperatures must be carefully balanced with elongation steps to ensure stability throughout the cycling process.
Following the review of molecular techniques, the summary shifts to the organization and querying of biological data using databases. The instructor explains fundamental database concepts such as normalization, indexing, and sharding, highlighting their advantages over simple file storage for managing large-scale genomic datasets. Specific bioinformatics resources like ENSEMBL, PubMed, PDB, and dbSNP are introduced, along with tools like BioMart that allow researchers to programmatically retrieve data without manual downloading. The lecture also covers the mathematical foundations of sequence analysis, including global versus local alignment strategies, scoring functions for matches and mismatches, and the distinction between linear and affine gap penalties which better model biological insertion-deletion events. Furthermore, the importance of substitution matrices like BLOSUM and PAM in evaluating protein similarity based on chemical properties is discussed.
The final portion of the summary addresses gene expression analysis through microarrays and RNA-seq data standards. Key topics include the necessity of normalization to correct for dye variations and experimental artifacts, statistical methods like t-tests and ANOVA for comparing groups while controlling for covariates, and multiple testing corrections such as Bonferroni and False Discovery Rate adjustments to manage Type I and Type II errors. The instructor also reviews file formats essential for bioinformatics workflows, including FASTQ for raw sequencing data, GFF for genomic features, VCF for variations, and PLINK/BEMAP for association studies. Additionally, the video covers clustering algorithms like single, complete, and average linkage, as well as distance metrics such as Manhattan, Euclidean, and Minkowski distances used to analyze expression profiles.
To conclude, the instructor presents example exam questions covering fundamental biological knowledge, such as the difference between pre-mRNA and mature mRNA regarding intron splicing, the function of tRNA in linking RNA sequences to amino acids, and the four main steps of a mass spectrometry workflow. Practical advice is offered on protein purification techniques and software development practices, including unit testing, regression testing, and test-driven development to ensure code reliability. The session ends with administrative details regarding the upcoming written exam format, registration logistics for different student groups, and an announcement that future data analysis streams will resume in April, encouraging students to focus their study efforts on the core theoretical concepts rather than introductory programming modules.
Read the full video transcript
back everyone if you're watching this on
youtube thank you for being here if
you're on twitch then also thank you for
still being here
um so let's just continue so lecture
nine
was about primer design so the first
part was about polymerase chain reaction
right so
[Music]
what do we need for pcr
so we need water nucleotides primers
template dna and a master student to do
it for us
but we also talked about what is a good
primer and when is a primary primer
right because primers need to be a
certain length and they need to have a
certain binding capacity by having like
an ac gt kind of uh so an a
atgc composition
and we also talked about advanced
primers so that you can do multiplex pcr
where you have or where you're not
amplifying using a single pair of primer
but you're using multiple pairs of
primers
we talked about universal primers and
semi-universal primers if you want to
amplify a piece of a virus but not just
from one strain but multiple strains
so then you can use things like
universal primers
i think the example that we did there
was
happy fe
something like that
but then here we also have guest mirrors
so if you don't have a dna sequence
available
then based on a known protein sequence
you can still design primers for your
animal
which doesn't have a genome sequence
available and you do this by back
translation right so you look at the
amino acid sequence and then you code
the amino acids back to their dna
equivalent
which of course is not perfect
so gusmar primers are generally longer
than standard primers to make them still
bind or have to still give them this the
binding properties that you need
so when we talk about pcr
pcr comes in three steps so know that
the first step is denaturation
where we heat up our sample to go from
double-stranded dna into single-stranded
dna and then the next step is lowering
the temperature a little bit so that the
primers can bind
so ahead this is generally like 90
degrees celsius annealing happens at
around
60
something degrees i think the the the
temperatures were mentioned in the
lecture
um and then had the primers bind to the
dna and then in the next step we
put the temperature up to like 70
degrees celsius and then the
the
the dna polymerase starts amplifying the
the dna and dna polymerase only starts
amplifying dna when when we have these
primers there so when when there's
double-stranded dna
so
we talked about pcr right so when we
have our gene of interest that we are
trying to pcr out then of course in the
first cycle we get like two copies four
copies eight copies and sixteen copies
um
let me just mute myself
i don't know what's going on with my
voice but um so hemp in in in pcr
we do an exponential amplification which
means that the number of cycles that we
use we can kind of estimate how much uh
product we're going to get
which is 2 to the power of the number of
cycles that we do
we also showed the first few cycles in
detail right because this is just a kind
of stylized picture and this is what we
would like to happen but of course this
is not how it exactly happens
but if you're interested in that then
please watch back lecture number nine
um so how we we also talked about the
length of the primer and the uniqueness
and the fact that when the primer gets
longer you need to have a higher melting
and annealing temperature right but
there's this trade-off because the
longer your primer the the higher the
chance that it's unique
but also the longer the primer the
higher the annealing temperature
and then you you end up in this
situation where there has to be enough
temperature difference between the
annealing step and the elongation step
and have we talked about things like
melting temperature so the melting
temperature is the temperature at which
half of the dna is single stranded and
half of the dna is still double stranded
so the melting temperature for
genomic dna generally tends to be at
around 93 degrees celsius but of course
for for primers this is much
lower because primers are a lot shorter
and we also talked about annealing
temperatures so the annealing
temperatures the temperature at which
the primer starts binding to the genomic
dna
have we talked about multiplex pcr
semi-universal primers guest mers and
many more and i also showed you in which
fields primer design skills are required
right so if you are ever going to do
real-time pcr or
studying population polymorphisms using
microsatellites or aflp markers and then
you you use primers
and we also talked about internal probe
design but then the most important
things to to remember about this whole
lectures is that if you design primers
you have to achieve the appropriate
hybridization specificity so it has to
be unique right you only want to amplify
one part of the genome and you have to
be able to do this stably
which means that you have to have enough
difference in the different temperatures
for the whole process to be able to kind
of cycle through these three different
temperature levels right because if the
primer is too long and your annealing
temperature is around 70 degrees celsius
right then the annealing temperature is
the same as the temperature used by the
polymerase and then stuff starts going
wrong because then then the stability
doesn't hold
and stability also has to do of course
with the number of ats versus gcs
because you want to have your primer to
bind very tightly to the dna and of
course this has to do with
the template dna and the ac gt content
of of the of the template dna as well
to remember right if you design primers
never design primers on repeated
sequences um because
if you design primers on repeated
sequences then of course they're not
going to be unique
so
primers cannot contain repeats
themselves either and so we want to get
rid of them and for that we can use a
tool like repeat masker right so you can
use repeat masker standard or you can
use it directly in ensemble so when you
export your sequence from an example you
have the option to mask your sequence
which means that it kind of blocks out
areas which are repeated across the
genome and it blocks out areas which are
of very low
low variability right if you have atc
atc atc atc then it will block out this
region so that you don't design primers
by accident to these kinds of sequences
so lecture 10 we talked about databases
so we talked about some terminology
about databases so what is
things like sql and these kinds of
things i try to explain to you guys why
we need databases because they have
advantages over just storing data on on
your hard drive and they allow things
like sharding which makes them much
faster and you can do indexing
and so now sharding means that you put
it on different sites so you have one
database which is not just local on one
position on the earth but which also has
for example the same database but then
in japan so that people in japan can
also use it and then you have indexing
so indexing means that you you look
through a column see what's there and
already start
anticipating on future queries
and by building an index so that you can
quickly find back stuff in the database
we talked about normalization of data
within a database
and
the organization of a database so that
you can search generally in different
ways and had that you have and we we had
an overview of all kinds of important
databases in
bioinformatics and biology
so you need to know that
what can we use ensemble for right what
what is in pubmed
but also that there are databases like
the pdb which focus entirely on proteins
that you have dbsnp which only focuses
on storing single nucleotide
polymorphisms in humans
and during this lecture we also talked
about biomart and i did a biomart
example using
our african
goat data i think
and hep biomart is a connector for r and
for other programming languages so that
you can automatically query
many of these important databases in
bioinformatics so that you don't have to
go to the database to the website and
start clicking on it and downloading
stuff no biomart allows you to write
code which will retrieve data for you
which of course makes your research much
more reproducible
so a little bit more about the
normalization right so normalization of
data means that instead of storing the
full student name in a single column and
you can break it down to a more granular
level saying that no every name has a
first name middle name last name right
and then that that's easier because now
we can see that all of the students have
the same middle name which of course is
not obvious from the data that we used
to have in in a single column
i told you about over decomposition so
when you start chopping up things that
actually should not be chopped up right
when you have a phone number then store
the phone number in a single column
don't start doing smart things like
storing region codes area codes and then
phone extensions
because these might change over time
plus that no one's interested in
no one's ever going to do a query give
me everyone who has this
phone extension right because that's not
a very good thing we also talked about
things which go wrong in in databases
right so for example duplicate data and
not so much duplicate data but the worst
thing is duplicate data which is named
differently right so if you have fifth
standard and fifth standard once you
write it with
the let up with the number five th and
the next time you write it with fifth
then of course um the database does not
know that these two things are the same
right so it is a duplicate data
but it's even worse because it's
duplicate data but it's also
inconsistent and this is the thing that
databases help you solve by using things
like foreign keys right so in this case
you would say no the standard is a
foreign key which points to another
table where all of the different
standards are described so we had a
bunch of slides about this and and you
don't have to know all of the normal
forms
but know that there are things like over
decomposition and that that generally
the the
database functions best when data is
more or less stored in a normalized
fashion
all right so in lecture 11 we talked
about sequence analysis so how we we
know that sequences change during the
course of evolution and so you have
point mutations where a single base pair
gets modified and but we also have
insertions and deletions and to make it
even worse we also have this sign of
xenologs and so genes which are
transferred from one bacteria to the
other but head generally these are
considered the three main forms of of
sequence variation uh during evolution
and then of course we talked about a
homology trick right so that since all
of all life on the planet more or less
comes from one single event where life
existed and then everything branched out
and we can use homology to kind of infer
the function of a protein right if we
know that a certain protein is
transporting oxygen in humans and then
if we find a protein in in mice
which has a very similar
which is a similar very similar sequence
and then we can also assume that this
protein and mouse probably also star uh
will transport oxygen right so that's
endomology trick works because it's a
single tree which kind of grows up from
from the early beginnings
when we talked about sequence analysis
we talked a great deal about sequence
alignment right so because that's kind
of the fundal fundamental algorithm in
in
bioinformatics and that there are two
major variants of alignment one is
global alignment and the other one is
local alignment so global alignment
tries to match the entire string to the
other string
so it is more likely to insert gaps at
the beginning
but when we um
i think something went wrong with this
picture
yeah this is this is wrong um please
please ignore this slide and and look at
the slide in the lecture i think i
copied the same one for
but her local alignment tries to find
the optimal substring while global
alignment tries to match the whole thing
so
yeah
yeah so this is this is just wrong i
copy-pasted the same thing more or less
twice but then look at the original
slide in lecture 11 and and know that
there's a difference between global
alignment and local alignment and local
alignment generally is used when you
have very short sequences which you want
to compare towards the genome while
global alignment is used when you have
complete viral genomes that you want to
align together
so
we talked about sequence analysis a lot
and um
when we talked about it we also talked
about scoring functions right so that to
compare if two sequences are similar we
need to kind of have a mathematical
definition of similarity right so the
the most basic definition
that we had or that we could come up
with is just the percentage of matches
so how many base pairs match and so you
get a plus one for each base pair that
matches and for every mismatch you you
get a minus one penalty um in a way
right so percentage of matches is just
seven out of twelve base pairs match
but you can also add um this this
plus one minus one system
and this plus one minus one system can
then be extended with a linear gap
penalty which means that when you open
up a gap in one sequence
then opening up a gap of two
compared to opening a gap of four the
gap of 4 is twice as expensive as the
gap of 2.
but since in biology we know that
insertions and deletions are very common
we nowadays almost always use an affine
gap penalty that means that you get that
you have a high penalty threshold for
opening a gap but then when you make the
gap bigger you don't put that much
penalty on there right so opening a gap
might be a score of minus one but then
going from a gap which is from one y to
a gap that is too wide you give a 0.1
penalty right and going from 2 to 3
again you get a 0.1 penalty so and this
is this this allows you to do much
better alignment when you use this
affine gap penalty so know what the
difference is between a linear gap
penalty and an affine gap penalty and so
the the main difference is that in a
linear gap penalty
you score for every
for every
x base pairs that you make the gap
bigger you get at plus x
penalty while in an affine gap penalty
you don't have that opening a gap is
expensive but extending a gap so making
the gap bigger is relatively cheap and i
want you guys to be able to calculate
the percentage of matches on dna and
protein level so when i give you one of
these
kind of amino acid codon wheels
you should be able to read the amino
acid codon wheel and if i give you two
dna sequences you should be able to say
well okay dna wise it matches 7 out of
12
but on protein level we see that there
is a 9 out of now a 4 out of 4 right
because of course we come in codons
so three letters of dna become one amino
acid
but be able to use these things so i i
think we did a small
example in the in the lecture as well
we also have to take care in dna
alignment that we score transitions and
transversions differently so
because when you when you
look at dna then it's
we had this
figure where
we showed that when you go from an a to
a t
that this is relatively common because
the
the chemical structure of the a
base pair looks very similar to the
chemical structure of the t-base pair
right so by using electromagnetic
radiation or or nuclear radiation um it
is very commonly uh it's a very common
occurrence for an a to change into a t
right but an a almost never changes to a
c or not almost never but when this
happens it's called a transversion and a
transversion is is very uncommon
and the same thing happens in proteins
because in proteins we have pro amino
acids which have very similar side
chains so head changing a a
lysine by a glycine is a very small
change because the because the side
chain is the same right and then we have
something which is the substitution
probability matrix where we talked about
the blossom matrix and the pump matrix
and these matrices they try to catch
this fact that
when we have a
substitution of one amino acid by a
different one so we have kind of a
mutation there um it tries to score um
some more heavily or it penalizes some
more heavily than others right if we
have a a
positively charged amino acid being
changed by a positively charged amino
acid then this is a relatively common
occurrence but a positive amino acid
which gets changed by a negative amino
acid that of course has a has a much
bigger penalty associated with it
know what blast is basic local
local alignment
and know about cluster w
so alignment multiple sequence alignment
where we try and align multiple
sequences together and cluster w
when we talk about cluster w we talk
about how to detect conserved residues
we can find conserved regions but we can
also find patterns in our amino acid
structure and when all of the amino
acids are for example positively charged
head then they don't all have to be the
same but then still there might be a
pattern saying that for the functioning
of this protein it is very important
that at this position you have a
positively charged
residue or positively charged amino acid
all right so lecture 12 was all about
gene expression analysis again we did
more or less the exact same thing as
what we showed in lecture one where we
looked at microarrays right so creating
oligo arrays had nowhere bioinformatics
is involved
you don't have to know all of the
different
things right but know that a tiff file
is just an image file with these dots on
the microarray which can be red green or
yellow
cell file is this
proprietary format from
alphymetrix
which stores
data about microarrays in kind of a
compressed way
i talked about normalization
why do we normalize during microarray
analysis and this is uh because the dyes
have various varying behavior right the
the green dye has a much higher dynamic
range than the red dye um but there's
also variation during the hybridization
like the the
surrounding temperature or the
surrounding humidity has a big influence
on how well your sample hybridizes to a
microarray and of course there's
variance in the manufacturing if i buy
an array now and i buy the same array in
like five years
then of course the the
the quality of the array might be might
be different right because the technique
gets better so the array that i do today
is not directly comparable to the array
that i do in like five years
so
to kind of get rid of these effects
right these are all effects that
introduce variants into the into the
sample and we want to get rid of that
and that is why we do use normalization
besides normalization of microarrays we
also almost always look at the log2
ratio of a microarray and this is
because of this varying behavior of the
dyes
where we want to have kind of a linear
scale saying that if i go from zero to
one um then this needs to be the same as
going from uh zero to minus one right so
it's a it's a it's a transformation
where we go and then we divide
the green dye intensity by the red diet
intensity and then we do the log 2 and
this is just to prevent the fact from
having
if you have 1 divided by 2 that is
different from 2 divided by 1.
and i think the slides actually in the
lecture explain it pretty well why you
want to use it
when we talked about gene expression i
showed you guys that you can do it using
t-tests but also you can do it using
anova test right so a t-test is really
nice when you have two groups
but when you have for example two
different groups two different tissues
and you have two different factors for
example low concentration of
of medicine and high concentration of
medicine then you are forced to do an
anova test right because an anova test
allows you to adjust for covariates um
it allows you to put up a model saying
that my intensity of the probe is
related to the condition in which the
probe was measured
plus the the
thing that i did to the sample
plus the type of mouse where the sample
was taken from right and you can't do
that with a t-test because the t-test
only compares two conditions so
condition a versus condition b and does
not allow you to to control for other
factors
here again i mentioned multiple testing
right so the type one error is calling a
gene significantly changed even if it's
just by chance
so type one errors can be avoided by
bonferroni correction and the type two
errors is when you say that a gene is
not significantly changed or not
significantly different between two
samples that you have
while it actually is and you can only
optimize for one of the two
so you can say i want to have a minimal
amount of type 1 errors but then of
course your type 2 errors go up
and you can say i want to minimize the
type 2 errors but then the type 1 errors
go up so that's the kind of trade-off
that you have to do
and had this one can be avoided by
bonferroni correction the type 2 error
is by mini hulkberg false discovery rate
adjustment
we also talked about gene ontology right
gene ontology is a
common
terminology we use to describe the
things like the cellular component where
the gene is found right so a gene can be
active in a nucleus a gene can be active
in the cytosol or it can be active
outside of the cell so that's exported
but also we have a common
nomenclature so a common terminology for
things like biological processes and
molecular function and this allows us to
do these over-representation tests right
imagine that i do a microarray analysis
and i find 50 genes which are different
between the two animals that i look at
then we can do these tests looking to
see if a certain cellular component is
over-represented in these 50 genes right
if all of these 50 genes are nuclear
genes then of course we we hypothesize
that there might be something going on
in the nucleus but if all of these 50
genes or 40 out of 50 genes are located
in the mitochondria head then we might
assume that no the mitochondria are the
thing where where it goes wrong or where
the animal has an issue
and the same thing for keg right using
keg we can actually
it provides these map of different
pathways in uh in different species and
we can actually overlay our gene
expression data onto a cac pathway to
see if everything in a pathway is
upregulated or if a whole pathway is
down regulated
based on the tissues that we're looking
at
again here we talked about similarity
right because we have to have a
mathematical different definition of
what is similar um and of course it's
different from when you look at when you
compare dna sequences to each other or
when you compare protein sequences to
each other when you compare expression
profiles to each other there are three
different distance measurements that you
can use right so the manhattan distance
is just the absolute difference between
the sample one and sample two and then
across all of the probes that you
measured the euclidean difference is
more or less the same but it's not the
absolute difference it's the difference
to the power of two
you add up all of these differences and
then you take the square root of the
total difference and then the minowski
distance is more or less the distance
generalization for this where you can
choose your own m factor right so an m
factor of two means that you have
euclidean distance but you can also have
an m factor of 3. and why do we
sometimes use minowski distance because
sometimes we want to put
more
weight on large differences
right because 0.1 to the power of 3 is
of course much less than 2 to the power
of 3. so by choosing a higher m factor
you're focusing more on extreme
differences compared to small changes
which are globally across
but three different distance
measurements to express
how similar or how different two
expression profiles are in animals or in
mice or in plants
again we want to build a tree right
because we want to see which things
belong together which things are not
belonging together and then we also
talked about this clustering um so head
there's a difference between single
linkage so if you have a group or a
group of two profiles and a other group
of also two profiles then the single
linkage is based upon the two most
similar elements within the groups if we
look at complete linkage then we look at
the two most dissimilar
elements and had the distance between
the two groups is then based on the two
most dissimilar elements and then we
have average linkage which is also
called op gma
and then we look at the distance between
two clusters is taken as the average of
all distances between pairs of objects x
in a and objects y and b and that is the
mean distance between the elements in
each of the two clusters so average
linkage is the best
but it is relatively expensive to
compute when you have literally hundreds
and hundreds and hundreds of elements
right if cluster one has a hundred
elements cluster two has a hundred
elements then you have to compare all of
them right so you have to compare one
versus a hundred
the second one versus a hundred the
third one versus a hundred so you you do
like a massive amount of comparison so
up gma is the most computationally
expensive method and that is why
sometimes people look at single single
linkage and complete linkage because it
is relatively cheap because you only
have to do one comparison
we also talked about where you can get
free microarray data
so go to gene expression omnibus if you
want to get free microarray data to work
on and write a scientific publication
without spending any money
and the same thing you can do at array
express
the massive advantage of array express
is that they have curated reannotated
archive data
which is a very high quality because
someone looked at it and made sure that
the sample that was submitted is really
the sample that people said that it was
and that's not the case for for gene
expression omnibus gene expression
omnibus anyone can upload data even me
so that means that there's no curation
going on
lecture 13 standards for analysis
very
interesting lecture i think because have
we talked about different
biological file formats like the comma
separated file but also fasta files so
sequencing files um
or
files holding sequence data then we have
fastq which is the standard output for
dna sequencers nowaday which contains
dna sequence data but also dna quality
data
we looked at the gff format which is the
the format for storing genomic features
and we have the vcf format for storing
variations relative to a reference
genome and we also looked at the bet map
format which is a very common format
when you do association analysis um so
it stores
variations on one side and other also
phenotypic measurements on the other
side so it's kind of the common file
format used in association analysis like
genome-wide association and and qtl
mapping
we talked about difference in testing
strategies so if you write code as a
bioinformatician then
use tests to test the code that you've
written right so a unit test means that
you test the smallest unit so you re
you've written a function so you you
throw all kinds of different input to
the function and then you see if what
the function gives you is actually
correct based on on what you wanted to
do with the function right
so regression testing is different
because regression testing means that
you take the code from someone else and
then just throw in data
see what comes out and then you start
modifying the code but you make sure
that every time that you make a
modification that you run the test and
make sure that for the same input still
the same output is produced
and then we also had some words about
test driven development where you where
you develop software using this
iterative approach where you say i i
want to add a new feature so i write a
test that tests the new feature the test
initially fails then i start writing
code
i then run the test again and if the
test succeed then i have successfully
implemented that feature and i continue
with adding a new feature read it so
it's it's writing tests and then writing
code to pass the test
we also discussed all kinds of different
types of documentation
so we talked about user documentation
and
documentation which is written
beforehand
we talked about
code documentation and so know that
there are different types of
documentation for different audiences
and so you're not only writing code for
yourself but you're also writing
documentation that belongs to the code
like a tutorial for people that will use
your code
but you also write things like function
descriptions so saying that this
function has five parameters
and these parameters
had the first parameter needs to be an
integer between 0 and 100
and so there's different types of
documentation for different groups of
stakeholders when you are writing
software
last lecture of last week i think that's
also one of the most fun lectures
because i always like doing it we talked
about citations why do we cite stuff in
in science and what's the use of it
we talked about things like web of
science so if it's not in web of science
it's not science we talked about google
scholar and research gate and things
like h indexes and i indexes
we talked about scientific reference
management and that you should do some
form of scientific reference manager
management using a reference manager
like endnote or mendeley and then i
showed you a difference between
distributed and centralized version
control right so that's that's it's not
directly related to literature
management but version control is
related to kind of software management
right because you you want to be able to
go back in time to re-run your analysis
and this has to all do with reproducible
research right so in theory lecture 14
should have been called reproducible
research instead of literature
but we talked about citations which are
there to make sure that
when you claim something that you point
to the guys that actually did the
research for it
have we talked about reference managers
which allow you to kind of easily
include references when needed
and version control is there so that
your code is also version so you can go
back in time
because of course code changes um
and
sometimes you need to re-run an analysis
as if it was 2017.
all right so with that
first exam question
example exam question
so if you throw the answers in chat then
we can go to the next one and we can all
go home early
so what is the difference between
pre-mrna and mature mrna
um i have a sound effect for that like
let's do the audio and then just do
crickets
so i'm just gonna continue this sound
effect until anyone answers the question
all right question uh answer number one
it's not spliced
hi by the way yeah hi shannon welcome to
the lecture um were you here the whole
time or did you just arrive um but
indeed indeed pre-mrna still has the
introns inside of the messenger rna just
forgot to say that's okay that's okay
it's good that you answer so
um yeah so mature mrna does not have
introns pre-mrna still contains the
introns
so head there's this process called
splicing which removes them so
that's entirely correct all right next
question next question
what is the function of trna
i'm just going to do crickets again
i have more sound effects we can do
birds
let's do birds for now and then uh
so anyone can answer like there should
be six people viewing it minus myself
and my moderator of course um
and then uh we can have an answer to uh
what what's the function of trna
i actually mentioned it during the
lecture i think so
it's the link between rna sequence and
the amino sequence of proteins yay very
good so trna is indeed it
it reads the code only in the messenger
rna and it it links it to an amino acid
so that indeed is the function of trna
so it's the uh
link between the rna sequence and the
very good
all right next question what are the
four steps in a mass spectrometry
workflow experiment
let me see um
everyone typing typing typing
good
just that sonic sin doesn't answer all
of them like misha come on you know this
misha like
get your head away from the olympics and
answer at least one of the exam example
exam questions
stop watching the olympics like
they will win the gold medals even with
you not watching them like
no i don't one gold in the pocket okay
okay so at least we want a gold medal
that's good that's good
all right so the four steps are of
course compound separation right because
we have a mixture so we need to separate
the compounds then we need to do
fragmentation and ionization
then it is
separation by mass over charge and then
it's detection i think i'm doing it
wrong
i don't have to know it fortunately you
guys do um so
that's that's kind of the way that it
works right so i already gave the
lecture so for the lecture i read up on
it
but yeah the four steps i think are
compound separation
fragmentation and ionization
separation using mass over charge and
then i think detection so that that's
kind of the four steps
by the way
i am relatively strict when it comes to
numbers
so if i ask you guys for
three things
and you write down four it is completely
wrong if you write down two then you are
maximumly allowed to score two out of
three points
but
it's not a guessing game
right so if i say um what are four
reasons to do x
and you write down six then it's
completely wrong
because i'm not gonna pick which four
are correct and which two are wrong or
the other way around right even if all
six are correct
then because i asked for four you did
not understand the question
extraction separation identification
quantification yeah that's what google
says but uh that's okay google can say
something else i think it's more or less
similar right let's just scroll back
quickly
so
metabolites compound separation
fragmentation and ionization separation
and detection
yeah see that's okay
all right um last example question i
think i had four
name two protein purification techniques
and describe how they
work so
that that
seems to be a difficult question i don't
think that i actually mentioned it here
because we did have protein purification
in the protein lecture
but i didn't i don't i don't think that
i mentioned this but uh
and of course you don't have to describe
now how they work because then you're
typing for like 15 minutes and uh but
these are the kind of questions that you
can expect right very basic questions
about the the the lectures that we had
um and generally i i like these two-fold
questions right so that you name them
and that you quickly describe how they
work
will the exam be oral or written um
what do you want it to be
because um i'd like to do a written exam
but in the pre-function studios are
known it says it has to be an oral exam
but then the question is because i
actually submitted a request to the exam
committee to have it changed from an
oral exam to a written exam and they
actually accepted that but it's still
listed in achnas as being an oral exam
while i'm actually so it's it's it's
probably going to be written
because i i think that written just
makes more sense
i'm fine with either all right good um
and i think i looked at acnes and i
think you are the only one who
registered for the first
period
for the second period there's actually
three people registered
so since you are the only one who
registered for the exam next week and
the other people all registered for the
makeup kind of exam or the second exam
date um in your case we can we can do
whatever you want if you say i just want
to have them
orally because they were through quicker
then that would be fine with me as well
but i i will think about it and i will
let you know because if two or three
other people of course still register
you can do a drawing question in an oral
exam
you actually can
because we're doing it via zoom
so sonic sin can just sit there make a
drawing and show it like that
right
and the question is still perfectly
valid like um draw
a platypus right
like it doesn't matter if it's written
down i can draw a puffer fish
that's a that's a that's a that's a
that's a challenge we can we can look at
that
um but anyway yeah i'm i'm still
thinking about it a little bit i think
legally i could do both um since i did
get the uh the the okay from the exam
committee to do it written
but on the other side you're the only
one who registered for the first date so
um we might just want to do it orally
then because that's going to be a lot
quicker and i can be a lot more flexible
right
if you
write it
down wrong then i have to kind of
say this is wrong but if you if you just
say it wrong i can kind of sit there and
do like
right so
we'll have to see i will let you know i
will let you know i will discuss also
with
the other people here
because when we do it orally i also need
to have a secondary
examinator there well i don't need that
for a written exam since i just have
your written exam
but i will i will let you know and i
will let you know before this weekend so
i will think about it i will discuss
with my colleagues here and then i will
let you know
tomorrow probably because i think people
can't register anymore for the first day
so
i think they can still register for the
second date but not for the first date
anymore anyway it doesn't matter too
much there will be an exam next week and
you will do perfectly fine
because already here you had three out
of four questions so that that's going
to be
going to be in the direction of like a
1.7 i think
but we'll see so
all right so um that was it for today
and for the whole lecture series
so
[Music]
i discussed with you guys all of the
different lectures that we had also
which lectures you should not focus on
learning so
don't don't spend too much time on the r
introduction lecture although it's an
important lecture because programming is
essential
there won't be any programming questions
on the exam
so that's it
so we're through
i think in total we did
50 hours of streaming 50 hours of
lecture
um so i want to thank everyone that
attended the
twitch streams um i want to thank you
guys for for being there thank you for
attending the course
um and like i said um i'll mail everyone
who registered for the exam with the
details as soon as i get the list still
didn't get the list should have gotten
the list already from the prefunctural
but they're not doing
and yeah good luck on the exam and
i am
i hope that everyone will pass so that
we don't have to have a
third exam date
and yeah i think thank you guys so much
for being here and i hope you guys
learned a lot of course if you have any
questions then
feel free to ask and
besides that
uh could we have that powerpoint that
acts as a study guide so you mean this
one
yeah i will i will upload it directly um
i didn't um upload it yet yeah no i will
do that directly um
good
all right then yeah thanks so much for
being here um i really enjoyed it um
it's nice to
stream it like or to be able to do it
like this at all um i miss the in-person
lectures i like the in-person lectures a
lot as well um but
i i think this is this is as good as we
can do with the current circumstances um
so thank you guys for for for being here
and spending
50 hours with me uh on on bioinformatics
and all of the different topics that we
discussed
um and uh
i will see xanax in at least next week
on the exam and the other
workers
students
i will probably see them on the second
exam date
and that that's it for now so unless
anyone wants to
get rid of some of their channel points
less denny bucks and have me make a
drawing
then i'm actually going to
close the stream early for today and
then
enjoy
learning for the exam and then enjoy the
weekend already
all right see you next week yes
and then uh
to all the other people that are still
watching bye bye and
we will see each other on stream
uh let me see let me see because i do
have another date
for the summer semester so streams will
continue or restart
um
let me see
so the data analysis using our course
will start on
21
april so the 21st of april there will be
a at least probably a stream unless i
have to do it in presents or if i can do
it in presents um but if it's going to
be online then 21st of january february
april 21st of april we will start the
course
so see you then and
i hope you enjoyed it i enjoyed it a lot
and uh thank you for being here