Video summary
Bioinformatics is defined as an interdisciplinary field that leverages computer science tools to address complex biological questions, distinguishing between raw data and actionable knowledge while operating in *in vivo*, *in vitro*, and *in silico* contexts. At the heart of this discipline are DNA sequences, which serve as the primary entry point for studies ranging from phylogenetic tracking of viruses like the coronavirus to whole genome shotgun sequencing. The curriculum explores various applications of DNA technology, including diagnostics, forensics, and biotechnology, while detailing specific sequencing methods such as Sanger sequencing and Next-Generation Sequencing (NGS). Students learn to interpret sequencing outputs, manage large computational files like FASTQ and BAM, and navigate the NGS workflow from sample preparation through alignment and variant calling. Furthermore, the course clarifies fundamental definitions, noting that a "gene" can refer to either a unit of inheritance or a complex genomic sequence containing introns and exons, and introduces transposable elements discovered by Barbara McClintock alongside regulatory mechanisms involving enhancers and insulators.
Beyond DNA, the course expands into RNA and protein biology, covering diverse RNA types such as mRNA, tRNA, rRNA, and regulatory RNAs like miRNA and lncRNA that play crucial roles in gene expression and splicing. The RNA-seq workflow is examined with a focus on measuring expression levels rather than identifying SNPs, utilizing tools to predict secondary structures based on free energy. Similarly, protein studies distinguish between apoproteins and holo-proteins, explaining how amino acids form chiral structures that progress from primary sequences to complex quaternary arrangements involving multiple polypeptides. Advanced topics include the use of machine learning in structure prediction via AlphaFold, separation techniques like 2D gel electrophoresis, and the evolutionary relationships between orthologs, paralogs, and xenologs resulting from gene duplication and speciation events. The study of metabolites is also introduced, classifying them as primary or secondary compounds and outlining mass spectrometry processes involving chromatography and fragmentation analysis.
Genetic analysis techniques are further explored through the examination of inheritance patterns, where traits are categorized as either Mendelian or complex polygenic characteristics influenced by additive mixing and dominance. The course explains how linkage determines the co-inheritance of close genes on a chromosome while recombination separates distant ones, utilizing two-point and three-point crosses for genetic mapping. To study these effects without the complications of heterozygosity, recombinant inbred lines are created through brother-sister mating to produce genetically identical populations, whereas back crosses lead to genomic imbalances that fail to capture dominance or additivity. Consequently, F2 crosses are highlighted as a superior method for investigating both additive and dominant genetic effects, despite their limitations regarding large associated regions. These concepts bridge into modern mapping strategies, comparing QTL mapping in structured populations with Genome-Wide Association Studies (GWAS) used in outbred human populations, where results are visualized using Manhattan plots to identify genomic regions controlling specific phenotypes.
The final segments of the course address the statistical rigor required in genetic analysis, emphasizing the importance of distinguishing between effect size and likelihood when comparing homozygote groups. Since researchers often conduct numerous statistical tests simultaneously, multiple testing corrections are essential to avoid Type I and Type II errors, such as misidentifying outliers caused by data entry mistakes like comma errors. The distinction between parametric and non-parametric statistics is taught based on the underlying data distribution, ensuring robust conclusions are drawn from box plots and histograms. Ultimately, the course synthesizes these diverse topics—from the molecular mechanics of DNA and RNA to the statistical frameworks of GWAS—to provide a comprehensive understanding of how bioinformatics integrates computational power with biological insight to solve real-world problems in health, agriculture, and evolutionary biology.
Read the full video transcript
welcome everyone um last lecture
bioinformatics um the summary so
if you've not watched any of the
previous videos this is going to be the
video for you because i'm going to
summarize everything in like two hours
um so it's gonna be good it's gonna be
good so
last stream last stream i'm it's like
puppy eyes it's such a shame such a
shame it's not gonna be the last stream
because the next um lecture series is
going to be online as well
um
i hope so i hope so it might be that
actually the university will force me to
do it in person um but we'll have to see
we'll have to see but i'm i'm very sad
about it like i like the like weekly
being here and and talking to myself and
reading chat and answering questions and
these kinds of things but um it's really
a shame that um we're done with this
series
so but you guys had 14 good lectures um
and uh i hope everyone does really well
on the exam so that uh
that we don't have to have a makeup exam
or something like that
anyway this is going to be the overview
right so the overview is just going to
be me talking through all of the
lectures and then all the way at the end
i have four
example exam questions they're actually
not example exam questions they're
actually real exam questions from
previous years
so just so that you guys get a feeling
of what kind of questions i i ask and
how i ask them
and i will be answering them so then you
can see what i think is important
but with that out of the way let's start
with
lecture one the introduction
um so in the introduction we talked a
lot about what bioinformatics is right
like bioinformatics is a discipline that
uses tools from computer science to
answer biological questions and then i
also gave you guys a whole bunch of this
kind of
definitions
but the thing is that if i would ask a
question like what is bioinformatics
right then i want you guys to at least
mention computer science and biological
questions right because those are the
two core elements in the in the
definition so you can write it down any
way that you want as long as it mentions
computer science or
information technology or some kind of
analogy for that
and of course the biological questions
need to be in there as well so that's
kind of the way that i kind of do the
slides right many of the slides they
have this kind of blue highlight
in definitions and that that's the
important part so for lecture one know
what an algorithm is know what it know
what data is know what knowledge is know
the difference between data and
knowledge and also know the difference
between like in vivo in vitro and in
silico because those things are are
pretty important for bioinformatics
we also talked about dna sequence right
that sequence is the fundamental data
type in bioinformatics that that's the
thing that started it all right so we
started doing like protein sequencing
and dna and rna sequencing at a certain
point and that's where the whole field
of bioinformatics comes from um yeah so
dna rna and protein sequences are more
or less the
the
reasons why the field of bioinformatics
exists because people needed to store it
in databases and then you need to
analyze it
and know that sequence is also the entry
point for many insidico studies right
if we think about the current
coronavirus situation then and it all
started with people sequencing the virus
figuring out that it's a cyber cough and
then making like phylogenetic trees and
tracking how
the virus
mutates across the world and spreads
across the world and had know what whole
genome shotgun sequencing is
and also know that sequence alignment is
one of the most fundamental algorithms
in bioinformatics so had the alignment
of two sequences against each other
and we spend a whole lecture on that so
we also went quite quickly through the
microarray workflow in lecture one um so
and know at which point you don't have
to be able to reproduce the whole thing
right like i don't expect for you to
like know point by point by point what
the microarray workflow is
but i want to be able to ask a question
like um
in which parts of microarray analysis is
a bioinformatician
involved right and then you could say
well biophysician is involved in
creating the arrays but also in data
storage data normalization and and
generally i will ask something like
give me
uh four steps or four things or three
things right
so
all right then lecture number two
phenotypes so phenotypes we talked about
qualitative properties and quantitative
properties so qualitative is something
that is um something like
it tastes good it smells bad
i don't like it or i do like it right so
it's something that you it that is
really hard to kind of put a number on
so and not measured with numerical
results
and then we have quantitative properties
um and and quantitative properties that
are properties that exist with a
magnitude or multitude
which means that they can be measured
using s i units
and we talked also about mendelian
traits and complex phenotypes so
mendelian traits are traits which are
caused by a single gene which causes the
difference in a phenotype while complex
phenotypes are phenotypes which are
controlled by many genes
so one example of a mendelian trait is
earwax
there's dry earwax and there's wet
earwax and there's a single gene in the
genome that controls if you have dry or
wet earwax right so it's a single
mutation in in a single gene
complex phenotypes are things like human
height
intelligence and all of these things
right that's determined by many many
different genes
we also talked about like this mixing
flowers thing um so
if a phenotype is additive right so if
the genetics underneath the phenotype is
additive then you get mixing
so that means that when we have a red
a red flower and a white flower right so
these are the gametes
then
you get this following mendelian
inheritance diagram
while if we have dominance right so one
of the phenotypes dominate
then we get a different proportion and
this is because one of the two and they
can't mix together
so if you have a red allele you will
always be red also be able to read these
kind of diagrams so um yeah i
might ask a question about a diagram so
i will show you a diagram and then ask
you is this a additive or a dominant
phenotype
furthermore we talked about the concept
of linkage right so because genes are
located on a chromosome
hey the closer they are together the
more often they are inherited together
so if they are very far apart then
there's a high chance that these two
phenotypes there will be a recombination
in between so homologous recombination
when the gametes are produced separating
the two phenotypes from each other
and so linkage is a very difficult
concept and i just want you guys to kind
of be able to tell me in your own words
what linkage is
and we also talked about two-point and
three-point crosses which are very much
the same um but you these are used to
determine if genes are linked or if
they're independent right so if they are
on the same chromosome and how close
they are on a chromosome and then we
have independent which means that gene
one is on chromosome one and gene two is
on a different chromosome for example
chromosome 11.
and the advantage of using a three-point
cross compared to a two-point cross is
that in a three-point cross you can also
infer the order of the chromosome right
so it allows you to kind of build a
genetic map where you can say well if we
start at the beginning of chromosome one
then we first see the phenotype for for
example broken wings and then we see the
the phenotype for eyes and then we see a
phenotype for antenna right so we can
determine the order of the genes on the
genome
and that is only possible when we use a
three-point cross because we can kind of
figure out if a is closer to b than it
is to c and these kinds of things
good
we talked a little bit about phenotypes
in lecture two as well so we talked
about visual and analysis like
box plots and histograms so be able to
kind of tell me things about box plots
or histograms right so a box plot
generally shows the median value
and then it shows the uh the quantiles
right so 50 of the data and then 95 of
the data generally in the vexes
um and then we have like things like uh
histograms that we talked about but that
probably really would be a question
about that
we talked about multiple testing so
definitely know the difference between a
type one and a type two error and we
also talked about descriptive statistics
right what is an outlier and how to deal
with outliers so you can you can
winsorize them away head generally
outliers are values which are very very
far apart from the distribution and they
can be caused by things like comma
failures where you write down the comma
wrong so instead of writing 3.0
you write
30.0 right so these kinds of things
happen
we also talked about things like
exploratory data analysis and to decide
which model to use on the data a little
bit
so if you have a really nice normal
distribution and of course you want to
go with parametric statistics
but if you have like a lot of outliers
in your data then it's probably better
to switch to non-parametric statistics
and a little bit about hypothesis
testing so
good and then in lecture three we talked
about dna right so we talked that dna is
used for diagnostics a lot used a lot in
biotechnology and forensic biology and
in virology right so dna is used to
catch criminals uh
find things like do you have the bracha
gene and have do you have a high chance
of developing breast cancer and these
kinds of things but also dna and dna
research is used a lot in biotechnology
right if we want to make like
algae produce fuel
then also we we look at the dna of these
algae and try to optimize them for
producing biofuels
virology i think that speaks for itself
so we also talked about the old more or
less classical ways of sequencing dna
data right so and we talked about
maximum gilbert sequencing pyro
sequencing summer sequencing and next
generation sequencing and what i want
from you guys is that you are able to
read these kinds of plots
right so here you see a plot which is a
maxim gilbert uh no this is a
pyro sequencing yeah i think this is
pyrosequencing and so you add the
nucleotides in order and then have when
the nucleotide gets incorporated you see
a little flash of light
and the height of the of the flash
determines how many
base pairs there were
but hey if if i would show you a figure
like this and i would say to you guys
like
this is the order in which the
nucleotides are added
then and i want you guys to be able to
say okay so this is then the resulting
sequence
the same thing for this um so maxum
giber sequencing where you where you use
like
where you use
kind of cutting enzymes which cleave at
different points so you have four
different cutting enzyme one which cuts
at a plus g one which cuts at a g one
which cuts at a g and one which cuts you
the c plus t which causes to which
causes the dna to fragment right and
then these fragments are brought up on
the gel
and then one of the things that i saw a
lot in recent years when we did the exam
is that people actually read it the
wrong way around
so the sequence here you read from
from the bottom to the top right because
here we see that it's c t a c g t a
and here you see c t a so hey you don't
read it and that's that's what often
goes wrong so people are able to kind of
figure out which base pair there was at
each of the different positions but of
course the first position is the lowest
one because the cutting enzyme cuts
at a certain point right so the smallest
fragment is the is the is the first base
pair
so remember that when you see a
sequencing gel
with maximum gearbear sequencing you
have to read it from the bottom to the
top
and that that goes wrong a lot so that's
a tip for you guys
we talked also about the workflow for
next generation sequencing and that you
do sample preparation then you do dna
sequencing generally
as a bioinformatician you're not
involved in that sample preparation is
done by a postdoc or by a phd or by a
master's student in the lab and the dna
sequencing is generally done by an
external company because like
at academia we almost never do our own
sequencing
but of course have what you get from the
company is these fast q files
and then we need to do all kinds of
steps before we can end up with a list
of our single nucleotide polymorphisms
so here we need to trim the reads which
means get rid of the
ends because this the quality of
sequencing drops so the more base pairs
i sequence the lower my confidence in
that the base pair is actually correctly
sequenced
so had a certain point we decided or in
the workflow you have re-trimming where
you say well i have my read my read is
150 base pairs long but i see that the
quality after like 110 drops off so then
i'm just going to cut the read there and
i'm going to ignore the last 40 base
pairs
so the re-trimming again yields a fast q
file same format as we had
and then we do alignment and alignment
is just taking the read scanning across
the genome see where it fits
and then here you get a bum file which
is this
kind of file format used for um for
next generation sequencing data
which is similar to the sum but then
binary
but
after alignment we have to handle
duplicates because in the sequencing
process we have generally a pcr step
when we do our sample preparation
but also we have optical duplicates
which are caused by how the machine
works
so
we have to remove duplicates which means
that if we have a read which is
starting at a certain position ending at
a certain position but we see the same
read over and over and over and over
again
then we we just ignore all the
duplicates and we just say no we had one
read at this position instead of having
like a hundred or a thousand right so
these optical duplicates are very
common so you have to remove those
then the next step is indel realignment
so indel realignment means that you use
known variation in the genome right
we've already sequenced hundreds and
hundreds and hundreds of humans probably
more in the order of hundreds of
thousands of humans by now
so if we do an alignment of a read
towards the human reference genome
then of course it might be that inside
of where the read fits
that there is a variant right so
a variant means that there's a single
base pair which is different in some
individuals
but that means that
we don't want to penalize the alignment
for having this variant
so that's what indel recalibration does
head looks at little insertions and
known deletions and then says well i'm
not going to penalize the read for this
because this is a known variant in the
human genome and of course we don't have
it just for humans but also for mice and
rats
and other model organisms
and then we have the base recalibration
step so the base recalibration step is
very similar to the indel recalibration
step
or the indel realignment step but the
base recalibration step is just looking
at single nucleotide polymorphism right
so it the indels are for kind of short
deletions and insertions and the base
recalibration is the same thing but now
for for single base pair variants and
then in the end generally what we do
with dna data is we don't look at the
whole genome that we have
but
imagine that we have a human then we
kind of want to summarize where the
human is different from the reference
sequence so we then do single nucleotide
polymorphism calling or snip and indel
calling
to find the regions or the yeah the
positions in the genome where our sample
is different from the reference
so also know that there are drawbacks
about doing next generation sequencing
right you need a lot of computing time
it's getting better and better because
tools become better and better of course
but in the end
there's a lot of computational time
involved in doing the analysis of dna
sec data you need a lot of hard drive
storage you need a lot of random access
memory and you need management of files
so hey you need to keep track of all of
these different files that are being
produced because of course in the whole
pipeline we start off with one file we
get from the company which turns into
two files three files four five four six
seven so it's like seven eight files
that you have in the end
and you need to manage those and those
need to be stored and you have to have
backups and these kinds of things
also know that there's a difference in
the definition of what a gene is
right in the previous lecture so in the
lecture when we talk about genes or as
in phenotypes right so units of
inheritance and but in molecular biology
and in sequencing a gene is not a unit
of inheritance a gene is a union of
genomic sequences encoding a coherent
set of potentially overlapping
functional products right because a gene
nowadays in biochemistry or in in in
molecular biology
we see a gene as having introns and
exons and a single gene can produce
different proteins or different variants
of the same protein
and so it's a much
more complex definition in molecular
biology than it is when we talk about
genetics in genetics a gene is very
basically a unit of inheritance so
generally it comes in two forms
you have a red gene and a white gene and
hey you get one of the gametes from your
father one of them from your mother and
these mix but of course in molecular
biology stuff becomes much more
complicated
because you have a gene which encodes
color and this color gene can have like
10 different variants right some of them
are white some of them are red some of
them are purple some of them are blue
and all of these things can mix and
match in an additive or in a dominant
way with each other so in molecular
biology
a gene is is a very um it's a fixed
definition and and it it's a it's a very
it's a very good definition but it's a
different definition than than what we
use in genetics so be aware of that
we also talked about transposable
elements there will definitely be a
question about transposable elements
remember that they are that they were
first described or they were discovered
by barbara mclintock
one of my favorite molecular biologists
ever so there might be a question about
that
but when we talk about transposable
elements so also called jumping genes
they come into two different classes so
you have retrotransposons and dna
transposons so the retrotransposons they
have an intermediate rna form
so they are more or less in the dna they
get more or less
transcribed into rna the rna gets then
built in into the dna as well and the
class ii transposons are dna transposon
so they don't have this intermediate
form
and then every one of these classes is
subdivided in two so you have anonym
autonomous retrotransposons and you have
autonomous dna transposons and
autonomous means that they don't need
anything to move so everything that they
need to move from one position in the
genome to another or to copy themselves
from one position to another
they carry with them
non-autonomous means that
it needs something from the host cell to
move from one position to the other one
right so it means that not all of the
proteins that it needs to jump around
are encoded on the transposable element
themselves
we also talked a lot about different
regulatory elements so just read through
them and know that there are different
types of regulatory elements like
insulators enhancers
you have tata boxes and and you have
like
metal sensing elements in the dna um but
i'm not going to ask in too much detail
about that i think that the transposable
elements i like that much more so
there's more likely to be a question
about transposable elements than there
is about regulatory elements
and also know the difference between a
mitochondria and a chloroplast know
their function hence the mitochondria
are the powerhouse of the cell which
means that they produce atp
chloroplasts are the same thing they are
also the powerhouse of the cell but then
in plants right so they they do
photosynthesis and produce atp for the
plant that way
lecture four rna so i think this is
generally the most
boring lecture for everyone
also for me
because there are so many different like
types of rna right so you have messenger
rna which comes in pre-messenger rna
called hn rna and then you have the
mature messenger rna called mrna we have
transfer rnas which transfer amino acids
so they form the the link between the
the messenger rna sequence and the
protein sequence by by having
well on one side they have the codon and
on the other side are on one side they
have the anticodon right which matches
the messenger rna and then you have the
amino acid which is attached to it in a
cloverleaf system
you have ribosomal rna
which is rna inside of the ribosomes
which helps the ribosome be able to
produce proteins we have small nuclear
rnas which are in the nucleosomes in the
in the nucleus which do things like
splicing
we have catalytic rnas like ribozymes
which
have a function themselves so it's
proteins generally have a biological
function but catalytic rna like
ribozymes they also have a catalytic
function so they are involved in
biological processes we have micrornas
which are there to do regulation of gene
expression so hentai generally are
binding messenger rna which then gets
degraded because the cell doesn't like
double-stranded rna
we have small interfering rnas which is
kind of a micro rna which is brought
into the cell by humans
or by micro injection
and we have non-coding rna which is
actually rna which does not code for a
protein
but we don't know exactly what it does
right so generally it's like the it's
like the micro rnas but then much longer
so long non-coding rnas or ncrna
so there are a lot of different types of
rna a lot of different types of
definition i don't really like to ask
very specific questions about it but i
do think that it's important that you
know that like rna is divided into all
kinds of different subgroups
again the workflow for
rna sequencing is the same as for dna
sequencing the only big difference is is
that you acquire your samples you
extract the rna instead of extracting
the dna and there is this additional
step where you do rna to dna reverse
transcription
and of course in the end because we do
rna sec we're not in we're generally not
interested in the snips so the
variations in the genome we are
generally want to do the extraction of
the expression levels at the end right
so instead of
saying that well at this position
my sample is different from the
reference genome in this case what you
are going to do is say well
i look at my gene of interest and i
count the number of reads that are there
and then i'm going to take the number of
reads in sample 1 and compare them to
the number of reads in sample 2 to see
if there's a difference in expression
level
so head the goal of rna sec is different
from the goal of dna sequencing in that
you want to get the expression of the
genome so the expression of the
different genes in the genome
while generally in dna sequencing you
want to look for variations in the
genome
we also looked at tools to predict
secondary structure of rna right so
generally you take the sequence you
annotate groups of secondary structures
and then this is all based on the lowest
free energy structure right so it tries
to fold the rna in such a way that
there's the least stress on the molecule
so here we have things like rna fold
which i think we had an example of but
there's also context fault and rna
shapes and there's a lot of different
tools
in the rna lecture we also
said that if you look at these
short rnas or if you look at these long
non-coding rnas right they generally
have like this modular structure which
means that
there's hey if you have a long rna
molecule then part of it can for example
bind rna or dna
part of it can bind proteins but also
parts of rna can be conformational
switches and these things they are
buildup modular right so you can have a
long non-coding rna which has two
conformational switches and a protein
binding domain
or you have a long non-coding rna which
has a dna binding domain a
conformational switch and then an rna
binding domain so and based on on which
which kind of structures we find in the
rna we can kind of figure out what the
function of this rna is right if an rna
has a protein binding domain right a
piece of the rna is predicted to bind
the protein then of course we can kind
of infer that this rna has something to
do with proteins so
but
it's a modular structure so these long
non-coding rnas they are modular so
they're build up different modules which
are more or less mixed and matched
together
all right so in lecture five we talked
about proteins
so here we have some nomenclature right
which i want you guys to know so an
amino acid is a single building block we
have a polypeptide which is a chain of
several amino acids and then we talk
about an apple protein which is one or
more polypeptides but not having the
cofactor
so for example the zinc molecule that is
needed to bind
the thing that it needs to bind or the
iron molecule to bind oxygen when we
think about hemoglobin
and then when we talk about proteins
right then we talk about apple proteins
with cofactors right so
hemoglobin is a protein and then when we
say the hemoglobin protein we mean the
four chains or the eight chains of
hemoglobin i think it has four
so it has four chains so there's four
apple proteins so four of these
polypeptides and then within these
polypeptides you have iron molecules
which bind oxygen right and then we talk
about a protein
so we also talked about hirality right
so the fact that
if you have an amino acid
right then almost all amino acids have
are chiral right because this molecule
in 3d right cannot be
put on top of this right it's the mirror
image right that is what chirality means
that you you have a molecule and then
you have the mirror image of the
molecule and these two although they
have the same structural formula
they do not have the same 3d structure
and because of that you can have
one of them being very toxic and the
other one being very beneficial right so
we also talked about that in nature
most of the amino acids are found in the
left form so the l form
and the d form is generally not seen or
it's generally not produced
but this the chirality itself in
proteins becomes a big issue when you do
like um chemical synthesis of um of of
medication right because when you do
chemical synthesis uh the chirality is
um egal right because we don't care
about the chira or the the the process
the chemical process that we use um uses
um like a plus b is c right but when c
is produced it's produced in both forms
so when you talk about amino acids
remember that they are chiral also
remember that there is one amino acid
which is not chiral and that is glycine
because glycine has an h as the r group
right the r group is the kind of side
chain which determines which amino acid
we're looking at and of course when r is
an h
then we are able to turn the molecule in
such a way that we we end up with the
mirror image so
glycine the smallest one so when the the
side chain is just a single hydrogen
molecule
then it is not chiral um you can draw it
and then try to uh try to do it
there's also these boxes actually um so
you have these
snappy atoms
snap ems or something like that they're
called and there you can just build
these amino acids right so you have c
molecules and you can stick in the
things um so if you if you are
interested in chirality and stuff then
then pick up one of these boxes of snap
snap or snap atoms i don't know exactly
what they're called
but then you can you can build these
atoms yourself which is really fun
so when we talked about proteins we
talked about the fact that you have the
primary sequence so the primary
structure is just the
amino acids in a row right so you have
glycine violin volume lysine isolyzine
so when we talk about the secondary
structure the secondary structure
and the primary structure of course and
this is what i pointed out in the
lecture um
is base the primary structure of
proteins is based on atomic bonds
and because of the fact that some amino
acids actually are able to form sulfur
bonds you can have
primary structures which are not just a
single line of more or less letters
right you i showed you guys i think
i showed you in the lecture two or three
more or less complex primary structures
where you have two polypeptide chains
which are connected together by a sulfur
sulfur bond um because of the
and and that is that is the difficulty
in primary structures for proteins is
that unlike dna and rna which is just a
single more or less straight line of
letters um in proteins the primary
structure
has already other
interactions and and so primary
structure is based on atomic bonding
secondary structure is based on um
hydrogen uh bridging right um so
and then we have the tertiary and the
quaternary structure so the tertiary
structure is based on more or less all
forces working on it and quaternary just
means that we take the whole protein so
the different
polypeptides
note that there are different
computational tools to predict protein
structure so there's up initial
prediction where you just take the
primary sequence and then try to predict
secondary tertiary and quaternary
structure
but we also have dedicated tools for
secondary structure prediction
because that is more or less something
that we can do very well but from the
primary structure determining the
tertiary structure is really hard
there's really good tools out there
which can actually predict if there will
be an alpha helix and if this alpha
helix will go through a membrane
because these things are very
are very common right so we know exactly
how transmembrane alpha helices look
like
there's thread and fold recognition and
homology modeling
so hen know that there are five
different more or less
schools of thought about how to predict
secondary tertiary and quaternary
structure of proteins from primary
structure
we also talked about the new alpha fold
from google
which is kind of using machine learning
to do it but again machine learning is
just the field of homology modeling
right because machines they look at all
kinds of examples and then they learn
how a protein folds based on the
examples but that of course is kind of a
type of homology because learning from
an example means that you use homology
we also talked about how you can
separate proteins right so
head there's two dgl electrophoresis
which allows us to separate protein
mixtures and we separate using two
different methods so the the standard
method is used for the y-axis or the the
the y component of the gel right so
that's the same for rna gels dna gls and
and protein gels and so here we separate
based on size using an electric charge
and then in the other
axis on the x-axis we separate using a
ph gradient so we start off with a very
low ph low ph of like two and then we
end up here with a high ph of like
14 right so water is like seven so in
the middle
so every protein comes with a charge and
that is because they have side chains
and the side chains they give this
protein an intrinsic charge which means
that a protein which has a positive
charge feels more at home in a negative
environment right and a negative
environment means that you have
an
abundance of hydrogen
so that means that you are then in a
positive ph
but i could be wrong right but had this
this second axis is based on the
isoelectric point and the isoelectric
point from a protein or
a protein isoelectric point is made
because of the fact that the protein has
side chains
we also talked about orthologs paralogs
in parallax out parallax center locks
and i want you guys to kind of know what
it is um and i
i hope i explained it well
but it was at the end of the lecture
after like a two and a half
hour stream
so if i didn't explain it properly um
in the in the lectures then do look it
up online
because it is important there will
definitely be a question about
what is the difference between an
ortholog and a paralog right so and this
has to do with the gene duplication
events and speciation events and so when
a species splits into two species or
when a gene duplicates itself across the
genome
and i hope i explained it well during
the lecture
but
if i didn't
then
please look it up because like i can't
explain everything perfectly because
if i've been streaming for two and a
half hours then sometimes the
quality of my thinking goes down
uh and besides that we have of course
xenolog so xenologs are more or less
pieces of
dna or proteins which are transferred
from one species to another right so
it's a horizontal gene transfer
mechanism and we we talked about like
four of them and one of them of course
is just cloning or genetic engineering
of bacteria but bacteria also exchange
dna with with other bacteria so they
make these little tubes and then they
just exchange parts of their dna with
each other to
increase survival for both of them
all right lecture six was about
metabolites so we talked about
endogenous and exo
endogenous metabolites and exogenous
metabolites
heso
know the difference between the two we
also talked about primary metabolites
and secondary metabolites so primary
metabolites mean that if you don't have
them you more or less die instantly
while secondary metabolites are
metabolized which you can go without
so
we also talked a lot about the mass
spectrometry workflow so
mass spectrometry is four different
steps
the first step is compound separation
which can be done using three different
techniques
two of them which are chromatography
techniques either using a liquid or a
gas as the mobile phase and then we also
have capillary electrophoresis which
again is very similar
to how we separate proteins and how we
separate dna by their size but here we
use electrophoresis using a very narrow
capillar
and in the capillar we kind of break
down or we we
we slow down big proteins because they
are big and and small proteins go
through relatively quickly
so after we've done the compound
separation in mass spectrometry we go to
fragmentation and ionization which means
that the the protein that or metabolite
that we're looking at gets fragmented
into little pieces and then each of
these little pieces gets ionized so they
get a charge put on them so this can be
two positive charge or three or four or
one positive charge and of course we do
this to be able to have the the the
thing fly through the mass spectrometry
right because
it needs to be charged to be attracted
or to be shot out
and then we have the separation of the
mass over charge right so we can do this
using a sector instrument or a
time-of-flight instrument and then we
have the detection so the detection part
is actually just generally a the the
charged molecule molecule flying against
the metal plate and then this is
detected using a computer
we also talked about keg so we talked
about
headed keg as pathway information and
that it's based on kind of a
input protein output
right so you have a metabolite and then
a protein working on that metabolite
transforming it into another metabolite
right so it has these kind of compounds
and reactions um hey so they have genes
and and proteins in there
but the main
selling point or the unique selling
point about keg is that it allows you to
reason
what type of metabolites an animal can
make and which type of metabolites the
animal cannot make we also have reactom
which is a different database which
is very similar right it also contains
pathway information it also has many
different organisms this one is open
source cac actually has a paid version
and also a free version
but the difference between keg and react
home is that
cag is very much based in more or less
chemistry right so we have a metabolite
and a protein working on a metabolite
transforming it into something else
while reactome is
more
holistic in a way right so they have a
pathway for rna
or for rna transcription or dna
duplication right so they're they're
they're pathways in reaction are very
similar to the pathways in keg but they
look at a slightly higher level so it's
not metabolite protein metabolite
it's it's it's more conceptual
we also talked about cytoscape
one of these open source tools that
allows you to visualize complex networks
and integrate with any type of attribute
data which of course means that it's
used a lot in bioinformatics to show
like large gene networks or large
protein net networks
but it's also used in social network
analysis which means things like
facebook and you can use cytoscape to
visualize your friends and who are their
friends and hey you can then use
different attributes so you can say well
everyone living in germany color them
green everyone living in in poland
calling them blue and all of these
things right so you can overlay all
types of different data on top of your
network and that is why cytoscape is
really useful
and we also have
it is also used a lot in the semantic
web
so have when websites are presented to
you
it's just plain text but you can use
html tags to html tags
to kind of give meaning to
parts of the text right so you can tell
for example the search engine saying
that
denny is a name right and it's a first
name while aaron's is a family name and
then the search engine starts to
understand what's going on and it can
build up kind of an internal network
saying that okay so
denny adams is a person and he works at
this department this department has
other persons working there and then it
can kind of form a more comprehensive
image on what is being displayed on a
website
and this is called the semantic web or
web 2.0 i think and nowadays people are
talking about web 4.0 i i got lost at
web 3.0 like for me it's all html cms
and javascript
but there's there's a
apparently a difference between the
world wide web now and like 10 years ago
i think the main difference is just that
the spying is increased a lot so
um then we had lecture number seven the
introduction into r and there will be no
questions about this on the exam because
this is just a lecture for you guys to
show you guys that you should if you
want to have a career in bioinformatics
you should definitely pick up at least
one data science programming language
so hey it's really good to learn
something like r or python
which are more or less the two main
languages that are being used in
bioinformatics
so but for you guys there will be no
questions about this on the exam so
that's good that means that you can just
skip the lecture when you are learning
for the exam
all right and then we went all the way
back right because now we discussed all
of the different biomolecular levels we
started off on the lowest level which is
the dna then the rna then the proteins
then the metabolites right and then we
started talk i started talking or the
lecture eight was about phenotypes and
how we do qtl mapping right so i talked
to you about the quantitative traits and
the i uh the ec also a little bit of
repeat of the first lecture
qualitative traits which are more or
less measured subjectively i showed you
this picture where we say that
quantitative trades are a subset of all
trades out there so all trades out there
are qualitative and quantitative
together but of course like quality
quantitative trades are a subset of the
qualitative trades and this subset is
growing right the more machines we build
the better we are in kind of um
expressing qualitative things into
quantitative units
right so an example of this would be uh
the the taste or the quality of wine
that used to be a very
qualitative trade right you would have a
panel of wine tasters everyone would
taste the glass and then they would
score the wine saying this is a good or
a bad wine but nowadays you just have a
robot that does that right so a couple
of drops of the wine get put in the
robot and the robot analyzes the
composition of the wine and then just
gives it its score
so
quantitative
is growing while qualitative is more or
less shrinking
and again i talk to you guys about
mendelian and complex phenotypes
so when we talk about phenotypes in qtl
mapping i taught you guys about the
crossover events right so that we have
meiosis 1 meiosis ii and that this whole
thing works or this whole thing is um
and that we can do things like
associate a region of the genome with a
certain phenotype or find a region of
the genome where a phenotype is more or
less controlled from
and that is only possible because we
have this chromosomal crossovers right
so that that in meiosis one has so what
we get is we get duplication of the of
the of the genomes that we have and then
we have the homologous chromosomes which
are more or less bound together and then
we have this
this crossover where parts of one
chromosome are exchanged to the other
one
right and then of course we have the
meiosis ii where now we go from having
two copies of each chromosome to having
only a single copy of the chromosome and
then these are called generally gametes
so and i also had two links there in the
in the lecture which are two
more or less
little movies where it is explained in
much more detail and graphically right
because a movie can show you guys how
this happens so that they align together
and that they then get swapped around
and that's very difficult to catch in a
in a picture
so when we talked about qtl mapping i
told you guys that you can only do
qtl so quantitative trade locus analysis
when you
use an experimental cross right so you
start off with for example two inbred
founders who get crossed together then
we get a generation which is called the
f1 generation and in the f1 generation
everyone has one chromosome from the
father one chromosome from the mother
right and there's no no crossover here
or no recombination because of the fact
that the parental line had two exactly
identical chromosomes so for the father
it had two exactly identical chromosomes
and for the mother the same thing so had
of course crossover occurred but this
crossover had no effect i think i even
made a little drawing during the lecture
showing how this works
but i want you guys to know the
advantages and disadvantages of the
different types of crosses that we
discussed and so
for example what a recombinant inbred
line is and so where you do this cross
between the two founders and then create
these funnels in which you within the
funnel you start brother sister mating
so to make sure that you get immortal
animals um which you can use forever and
ever um well they're not immortal but
they're like clones right so a single
real line so recombinant in red line so
one of these lines
you can mate a male with a female and
then the children of these will be
exactly identical
but there are of course
problems there because if you have a
recombinant in red line then there's
only two states so either being a a or
bb
so there are no heterozygote animals
within the population and because there
are no heterozygote animals in a
recombinant inbred line you cannot
estimate things like additive and
dominance you can only see that there's
a difference between the two homozygote
groups but you get no information about
the heterozygotes in the middle
the back cross is more or less the same
thing
but the back cross is really quick to
make because you cross two inbred
animals you get an f1 generation and
then this f1 generation is crossed back
to one of the two
parentals and then have we have the
we have the advantage that it's really
really quick to do because you only need
two generations but the problem with the
back crosses is that when you do the
association you get large parts of the
genome which are associated
and you have this imbalance between
having only 25 percent aaa and 75 bb
and again because individuals are only
a a or a b you get no information about
dominance and additivity
f2 cross more or less
solves these things
so it has the disadvantage of still
having like large regions but it allows
you to investigate additive and dominant
effect
so qtl mapping gwas we talked about the
difference we talked about how they are
very similar right so there's there's
they're both methods to find regions of
the genome which control
genes are which control phenotypes
right and
the the differences is that in qtl
mapping you are able to map between the
markers because of the fact that you
have a structured population while in a
g wash you just have an outbred
population generally
of humans and there of course you cannot
know what is between the markers but in
a in an f2 for example you can map in
between the markers and of course
there's another big difference and
that's the way how these results are
displayed so in a qtl you have a smooth
line plot across the chromosome and in a
genome-wide association you generally
have the results presented to you as a
manhattan plot
so had during the lecture we saw
examples of that
i talked also about effect size versus
likelihood so that the effect size is
the the difference between the aaa and
the bb group and that the likelihood is
the statistical test when you compare
individuals having a a versus
individuals having bb
and of course here we have to also think
about multiple testing
but that came back i think in another
lecture as well so multiple testing is
of course the issues that when you do a
lot of statistical tests
you have to kind of compensate for the
fact that you did a lot of tests
good
so i've been talking for around an hour
so we'll do a quick break and then we
will do the
remaining five lectures and then go to
the um four example questions
so that you guys have an idea of what i
ask and when
um so
let me set up not the audio but the
music um so yeah i'll be back in like 10
minutes and then we'll just continue
with uh discussing the different
lectures so five lectures left and then
then we're done so it's going to be a
very short lecture so i will see you
guys in around 10 minutes if you're
watching this on youtube then um
probably see you tomorrow so bye bye for
now