Video summary
Mike Sager from Family Tree DNA presented an in-depth overview of his extensive project, the Tree of Mankind, which aims to map the evolutionary history and branching patterns of humanity through genetic data. Unlike mitochondrial DNA trees, this project focuses primarily on the Y chromosome, which is passed unchanged from father to son, allowing researchers to trace direct paternal lines back to a common root for all males. The presentation explained that building this tree relies heavily on Next-Generation Sequencing (NGS) data, specifically the Big Y test, which sequences large portions of the Y chromosome to identify thousands of specific mutations or Single Nucleotide Polymorphisms (SNPs). These SNPs serve as unique markers that define haplogroups and sub-haplogroups, with Family Tree DNA playing a central role in naming and standardizing these variants to prevent the confusion caused by inconsistent naming conventions from different laboratories.
The speaker detailed the historical evolution of the tree, noting how it transitioned from an academic consortium effort in the early 2000s to a rapidly expanding community-driven project led by Family Tree DNA. Initially, the tree was static and updated infrequently, but with the advent of high-throughput sequencing, the number of known variants has grown exponentially, adding thousands of new branches every few days. A significant portion of the talk focused on the interpretation of "phylogenetic equivalents," which are groups of mutations that occur in a specific order but whose exact sequence is unknown until more testing occurs. Mike illustrated how new samples often reveal previously hidden relationships, such as when a single individual splits a major branch or when ancient lineages are discovered outside their expected geographic regions, fundamentally altering the understanding of human migration and history.
One of the most compelling examples discussed was the discovery of deep splits within haplogroup D, an Asian lineage that had long been thought to be absent from Africa and Europe. Through extensive testing of individuals with West African ancestry brought over during the slave trade, researchers identified multiple distinct lineages that split from the main root approximately 65,000 years ago. Further analysis revealed even older splits occurring around 25,000 years ago, demonstrating that ancient populations once lived in these regions and left descendants who are still alive today. The presentation highlighted the collaborative nature of this work, where data from various sources, including academic papers and citizen science projects, is integrated to refine the tree's structure, though Family Tree DNA maintains strict privacy standards by not publishing individual names or surnames on the public tree.
In conclusion, Mike Sager emphasized that while much of the data collection and visualization is automated, the critical analysis and decision-making regarding how branches are added and named remain a manual process to ensure accuracy. He addressed the limitations of current predictive testing for STR results, explaining that Family Tree DNA has chosen to be conservative in assigning haplogroups to avoid misidentifying customers who might order expensive follow-up tests based on incorrect predictions. Looking forward, the speaker expressed interest in integrating ancient genome data from projects like the 1000 Genomes Project into the tree, which would help calibrate mutation clocks and provide a clearer picture of human evolution over tens of thousands of years. The Tree of Mankind continues to grow as a dynamic resource that combines cutting-edge genetics with historical context, offering unprecedented insights into how all humans are related through a shared paternal ancestry.
Read the full video transcript
ok ladies and gentlemen make things me
great pleasure to introduce our next
speaker someone whose appearance here
has been much anticipated by the genetic
genealogy community and that is Mike
Seder of from Family Tree DNA now Mike
might never expected to be in charge of
the tree of mankind a passion project
for Mike and he has done such a
wonderful job putting together at the
tree of mankind
and he totally of all the people on the
planet has the greatest overview of how
mankind has evolved and branched into
different branches over the course of
time so I'm really looking for this
presentation please welcome Mike singer
alright thank you very much ok so today
I want to talk a variety of topics the
tree of mankind not necessarily the
mtDNA tree but the tree in general maybe
some tips and tricks to interpreting the
tree how ft DNA is building the tree and
a little bit of history behind it and I
want to touch on some of the more
notable samples that mtDNA has produced
recently so I'm gonna try and stay away
from a real basic stalk but I do want to
cover a couple things first so the y
chromosome is passed down unchanged from
father to son because of this we are
able to trace back an entire line all
the way back up to essentially the root
of mankind so what we're able to do is
if no other data exists that I could
take every male in this room sequence
them and then build a tree
and tell you exactly how everybody is
related to everybody else about how far
back in time and about how closely you
are related so yDNA is very unique in
that aspect so just a little bit about
it the y-chromosome is currently about
57 million base pairs long which may
seem like a lot but it's actually the
second smallest chromosome behind comes
I'm 21 again passed down father to son
when we do wide sequencing we have to
have something to compare sequences
against so we have what is called a
reference sequence there is actually
nothing special or unique about it it's
just a universally accepted sequence
which everybody uses to compare against
these are updated regularly because we
don't actually know the entire sequence
of the white chromosome or any
chromosome for that matter there are
regions that are difficult to access
very repetitive stuff like that so new
references are being updated all the
time the latest one was in December of
2013 one before it was 2009 so we may
have a new one coming up soon that
actually doesn't have much to do with
genealogy because again the advancements
on the reference are usually going to be
more fringe elements that we aren't
using for genealogy just a basic
structure of the white chromosome here
that this grey part is a very large
portion which basically we don't know
what's in the reference sequence this
these parts in blue up here are called
the euchromatin
and this is basically where all of the
the great stuff for genetic genealogy
and the tree building comes from so it's
actually about half of the chromosome
that we are using so snip names the
information contained here the
there's the first part when you see
something like our m269 the first letter
is always a prefix that is basically
whoever named or discovered a particular
mutation in these examples B Y is Family
Tree DNA that's what we use for our snip
discoveries we did switch to F T when
Big Y 700 came around and we basically
did this because we're up to about a
quarter million variants and getting a
have a group of jby two three eight
seven one three it can be a little bit
cumbersome so this next part is just a
sequential numbering be y 10,000 is
simply the 10,000th BY marker that was
named this here is the chromosomal
location again the Y chromosome is about
57 million base pairs long and so B Y
10,000 simply exists at position 13
million 350 4622 so it's basically just
the address of the mutation and then it
also describes what the mutation
actually is the ancestral base is a T
and it mutates to an a this is just a
sample of the naming entities so you may
see a whole bunch of prefixes B y FG see
why this is a non exhaustive list you
can find this on on I SOG there's about
half of them but there's anything from
academics to citizen scientists looking
at Big Y or other NGS data and then
consumer genealogy companies like FG DNA
and FTC so
so typically the tree is built around
the Big Y or next-gen sequencing data
after we sequence a sample it goes
through automated variant calling and
imagining new snips are identified at
this stage and Ft DNA names them usually
within about 24 hours of a sample
posting and then we post those nicknames
online we can talk about that quite a
bit but we are trying to get names out
into the public because there used to be
a lot bigger problem in the paths of
double and triple naming a lot of our
snips were being renamed by other
companies other interpretive companies
or by other analysts and it causes a lot
of confusion so we just wanted to try
and cut that down still a bit of a
problem today but not nearly as bad
samples have been placed onto the FT DNA
haplotype
before I have really had a chance to
look at them and then at that time
matching variants and equivalent
breakers are identified and tree
structure is reviewed and I'll go on to
bet about that here in a minute so some
people seem to have a bit of difficulty
interpreting the tree one of the things
I like to suggest is to view mutations
as actual men's names because in all
actuality in all practicality that's
exactly what it is be wide ten thousand
had to have existed and occurred in one
man in the past and one man only so you
can just view all of this clunky
information as say mark an example of
that in use so here we know that Bob is
a descendant of Tom and then we know
that Tom and Bob are descendants of all
these men here but you'll see all these
men here are stacked up into the same
line or the same block
these are called Philo equivalents and
that simply means that we do not know
the order in which they occurred
so everybody tested to date either has
all of them or none of them so now say a
tester comes and he tests negative for
bill Charles and Arthur so what does
that tell us that tells us that we know
bill Charles and Arthur occurred more
recently while everybody else occurred
first more distantly so then we can
update the tree and we know that bill
Charles and Arthur down here in this
block and and the others remain up top
and so more testing may shed light onto
the order of these but that's just
typically how equivalent breakers are
dealt with so a little bit of history on
the tree itself in the 90s and early
2000s there was a lot of different
researchers academics pursuing their own
trees and phylogeny naming snips their
own and that led to a whole bunch of
confusion within the academic community
so a lot of them came together and
decided to form what's called the
y-chromosome consortium in 2002 and this
was in an effort to unify the tree and
make something more universally accepted
so this first tree contained 245
variants and 153 branches haplogroups a
through are are defined mr. t over there
you're not known about yet half the
group s has not yet discovered and
actually there was a little bit of a
caveat Kaplan group T is known about but
it simply resides here at the root of
half a group k so typically a new branch
is not added to the tree unless it is
viewed in two men thus the term
haplogroup anything else if it's just
observed in one person it's viewed as a
singleton however
this initial tree they just wanted to
gain some structure so that where
singleton branches added and so have a
group T actually does exist here but as
a singleton so it was not named yet okay
all branches today maintain what is
called monophyly which is which simply
means that they all share a common
ancestor so if you look here and have a
group II you'll see that everybody in E
has the same common ancestor and that is
e the root of e the only exception to
that is the oldest haplogroup which is
haplogroup a and because of the nature
of what we keep discovering older and
older branches many people in haplogroup
a are more closely related to people in
haplogroup r than they are to other
people in haplogroup a so that's kind of
an interesting way to look at it
haplogroup a is as old as everything
else down here combined and yet we know
the least about it and haplogroup a
double zero which is the oldest Y
chromosome that we know is even older
and we know even less about it so there
there also several interior nodes here
that would they would later go on to be
discovered but not by the YCC you'll see
here under K there are many branches
there's P o in own and several singleton
branches so the internal structure
grouping these is not yet known
haplogroup naming in in 2002 they also
came up with a system for nomenclature
for the tree two main types and those
two main types are still in use today
the haplogroup by lineage which is what
you're familiar with with our one b e1
a1 then that's still used by I SOG if
you go there website and it's used by F
T DNA
but only internally it's it's useful for
coating the structure of the tree and
moving things around but it's it's a
little bit more difficult for actually
communicating lineages because it can
change over time if we use that for some
of the people that we have in haplit
group J we'd have haplogroups that are
30 characters long and any change to the
structure of the tree it'll change year
after year the most commonly used one is
this hat nomenclature permutation and so
that's what you're more familiar with
something like our m269 G in 201 those
will never change regardless of how the
tree structure changes are in 269 will
always be armed 269 so there's less
ambiguity and it's a little bit easier
for the community to use I just thought
this was interesting in in this example
that they put out they show how to do a
branch split and they showed here and
have a group H what happens if in 52 and
m69 split off well that actually did
happen not too long ago but it was in
the reverse order chem 69 is proven to
be the parent of m-52 in 2003 two of the
authors came up with a slight revision
to it or an update to the YCC tree and
in it they expressed their desire to
resolve multi-fork asians into
bifurcations and that's just simply a
fancy way of saying that they are
looking for this internal structure here
to add branches to places where branches
could exist they just haven't discovered
them one example that they did find here
was grouping haplogroups in and oh so
they they they notated that one and then
they notated R - R - was discovering
that's a primary
the Arabic haplogroup actually quite
rare I'm still missing our haplogroups s
and t and at this time the Kois an and
ethiopian white chromosomes were
believed to be the oldest but it would
still be many years before we find y
chromosomes that are hundreds of
thousands of years older than or tens of
thousands of years older than in these
the next update wouldn't come for
another five years and basically they
doubled the size of the tree
they went from 153 branches to 311
they went from 243 bearings to 599 and
finally haplogroups s and T are
documented and today no other new
haplogroups have been discovered they
did add a couple of intermediary branch
of CF and IJ is finally joined but they
would they would never find any of the
resulting structure that consumer
genealogy would later discoveries such
as I and J will wind up being grouped
with K and an H will be grouped with ijk
and there are several other interior
nodes that actually genealogy will solve
rather than academia so the YCC ceases
to update in 2005 ft DNA released its
first tree which was basically a mirror
of this ycc tree
shortly after then within a couple
months I saw updated a tree and they
what their goal is was to kind of
centralize the effort to make a standard
where the YCC is no longer active in it
archived ice on trees are readily
available from 2006 all the way through
today it's still being updated and at
this point is when basically the
genealogy community
begins to outpace academia I'm not going
to spend a whole lot of time on this
this is just some some graphics that
show the growth of the tree over time
and see it starts without a hundred
snips a year 200 snips a year and we
started getting to about a thousand a
year May 2015 is probably about when I
started a little bit before then I did
and then we started really with the
introduction of Big Y and other next-gen
sequencing we're able to really grow the
tree at a rapid pace and if we had today
on here it'd be somewhere around 215 to
220 thousand variants on the tree as a
graphic that shows the FT DNA branches
over time as you can see were just over
25 I think if today was on there would
be right around 27,000 branches so what
the YCC was able to do in five years
jumping from 153 to 311 we're doing that
easily weekly really about every every
other day or every couple days where
we're adding that same kind of growth so
the Big Y 700 was a new product that we
came out there our goal was to sequence
as much of the Y chromosome as we
possibly could for a validation of this
we chose 88 samples most of these were
we picked there's many different
haplogroups as we could a double-0
are but then we also chose 11 samples
from half a group jay-z s 17 16 which is
Bennett Greenspan's haplogroup mtDNA
estimates this to be about a thousand to
fifteen hundred years old so what we
wanted to do is we wanted to see what
Big Y 700 uncover
that 500 could not so we use snips that
were called in a minimum of two samples
and then we used amanda was one branch
upstream at B y 101 as an out-group to
eliminate variants above that level so
what the original Big Y showed we found
16 variants and 7 branching points with
the old one big wide 700 retained all 16
but it also in covered an additional 8
new variants which is coincidentally
we're covering almost exactly 50% more
the y chromosome and we got 50% more
snips in just this branch 6 of these 8
proved to be equivalent to known
branches and to prove to being new
branching points one of these branching
points if you look under zs 1724 we have
s 5 s 6 and then b y 1 7 1 904 so we had
three known lineages behind this
we knew however that s5 and s6 were each
other's true closest matches we know
that because that hadn't actually been
it and his son and and these are there's
by1 seventh one group over here is i
believe their second or third cousins
but the old big y couldn't find a
variant that would group them but Big Y
700 did uncover a variant that we called
ft1 so now they have their own branch on
the wide tree similarly under Xia 1716
what we thought were three lineages we
had B Y 1 7 0 0 1 3
Ziya 1707 and then this lone man out
here s 11 well we actually uncovered a
variant that groups s lemon with this be
Y 1 7 0 0 3 group so that is a new
intermediary branch that we call ft 2
so I just had to include this graphic I
think it's one of the the most beautiful
graphics of the tree that that I've ever
seen gives you a real nice perspective
of saturation in certain haplit and
others hamlin group R is young and in
the in the white tree it's one of the
youngest haplogroups that there are yet
obviously it's the most well defined
well tested and here a zero as old as
everything else here barely gets an
honorable mention up here but you see
that our I and J really dominate the at
least that the growth that has been
driven by genetic genealogy and this is
the actual tree structure the FT DNA
tree structure I just love that graphic
so I'm going to give an example of how
the FT DNA tree is built off of everyday
results so I took an example this is
from haplogroup J a branch called z18
271 there were a wave of big Y's that
came in from this group and so what I do
is I I take the person's novel variants
and other variants that are unique to
just that group that have at least one
ancestral value in this group so
currently there are seven people at this
group we call them a through G and then
we have these shared variants here these
are the variants ref just means that
they're reference are their ancestral
they don't have anything there and then
this T these green boxes indicate that
there is a variant present so does a
simple reordering shows that there are
three branches within this
that could be added to the tree first we
have this single variant here that is
shared between samples enf simple enough
then we have another block here that is
shared between samples D B and G you'll
notice here that sample B has no
coverage that that's where this end
means that he has no coverage for these
snips however we know that because he
forms a branch with D that he has to
have these mutations so it's not
necessary for him to have coverage here
in all in actuality well this is D and G
have big wide 700s and sample B was from
big y500 so here is the third branch now
for a little bit of a trick eNOS added
into it this down here is a thing that I
run for the total number of all calls
are the total number of snips that are
viewed in the EPI DNA database and I'm
finding three people with this variant
well I'm only showing two here so where
is this other snip well it's a card in
the tubes the 18 to 7 1 min that we've
mentioned and then it also occurs in
this branch which is very close to here
so what exactly does that mean
if we go and look on our tree these men
are currently placed up here and they're
negative for everything downstream but
they share a mutation with this man so
how could that be another way to view
this is if I search this mutation and
all of these men this everything in
highlighted in yellow here is everybody
downstream of z18 290 you'll see that
nobody has coverage except for the
single big wide 700 right
so we know from this that everybody in
this group has to have this mutation so
being that we know that we know that
this nip is actually upstream of z18 290
and so when we add that that to our tree
we have a new big wife 700 branching
point downstream of Z 18 to 7 1 and then
the two branches that we referenced in
the discussion are down here so there
were actually four branches of the tree
that were added so there is a bit bit of
trickiness comparing Big Y seven hundred
to five hundred but it's really
uncovering a lot of neat branches like
this and this is just a low-level one
we're finding a lot of higher-level ones
that are proving to be quite interesting
this is just a block tree view of the of
the same branch that was added okay so I
want to talk a little bit about some of
the what I think are the more notable
samples that FDA has produced within the
last year year and a half or so so when
somebody takes an STR tests with ft DNA
we give them a predicted haplogroup we
don't predict very far down we're very
conservative with them but em2 is
actually one of the places that we
predict to what we had a man come in
from Saudi Arabia that that actually
split em2 it's the first time I'd seen
somebody split a branch that we actually
predict to its that that high up in the
tree this is almost a primarily West
African haplogroup you can see some of
the countries here that are most common
in our database and theorize for having
east to west
pension and so who breaks it of course a
man from east of West Africa yeah you'll
you'll know from our previous example
he's negative for all of these mutations
and then there's about a couple hundred
mutations that are still blocked up in
this m2 block early r1b splitters there
was a neat result that came from from
France actually late 2018 a man split
pea 3:10 and l1 51 now these are way
high up in the tree these are the
parents of p3 12 and you 106 so that was
pretty neat we have nearly 18,000 big
Y's and tens of thousands if not
hundreds of thousands of STRs downstream
of these branches but also in the end of
this year we found another man who fits
into this exact same line however they
do not form a branch to each other so we
actually uncovered a third line so
another way to view this is we have
three distinct lineages from P 310 one
with basically our entire database
lineage B has only one person and when
you see has only one person so I mean in
such a heavily tested group it's really
hard to wrap your mind around how
amazing that is
and and what's still left out there to
be found
here's another recent split in what I
saw calls have a group a one a this is
another West African branch found very
predominantly in places like Mali it's
actually very common in the u.s. it was
brought over in the slave trade mtDNA
estimates this branch should be about
51,000 years old with the tmrc a of
eighteen point four and who was at this
split this branch a Saudi Arabian so
when I was talking about the STRs and we
run for unpredictable people we run what
is called the snip occurrence program
which is a small number of snips to give
somebody an actual haplogroup so I've
done tens of thousands of these and in
2018 a customer came back and I was
scoring his backbone results and he came
back as CT star so a star where an
asterisk simply means that he's negative
for everything downstream so he's
negative for that this is a
visualization of the backbone of the
tree so he's positive for Beatty CT
negative for everything here and F is
the parent of ijkl R so we're gonna run
a Big Y on this person of course with
his permission
but there's only ten possibilities that
could come out of such a result and
every one of them are going to be
essentially groundbreaking and in terms
of the tree possible things he could be
a true CT star and form a new branch
down here that is probably a hundred and
forty thousand years old he could he
could fall down de it's a test that it's
a branch that we don't test for in the
backbone because nobody ever falls on
just that branch there either diri CF
another branch that we don't test for
cuz nobody belongs there and he could
also split any of these groups here
because all of them have equivalents and
we just test for one so you can split DC
e CT de see if doesn't matter which one
of these occurs that's going to be quite
interesting and Big Y is about to figure
that out
and what he actually did was he split
the root of haplin group d mtDNA is
estimated that this split is about sixty
five thousand years ago so that means a
man lived sixty-five thousand years ago
has descendants alive today that the
community had no idea about the fact
that that people are still out there
like that even an under test that how
the groups is still just it's
mind-blowing to me so he was found to be
derived for only thirteen out of about
250 snips at the root of happen or be
the participant could trace his lineage
back to Al wash and Saudi Arabia along
the northwestern coast of the Red Sea
this between Egypt and Jordan so we ran
a big why on him and then we also asked
him for his most distant known paternal
relatives so that we could run a big why
on him unfortunately for him he was not
able to trace his mail line back very
far and the best we could get was a
first cousin so we ran him we'd found
only one snip difference but then we
added this branch to our tree and called
it D F T seventy-five or it could also
be referred to as d2 and so this branch
formed about sixty-five thousand years
ago but everybody and it has a common
ancestor about a hundred years
haplogroup D is an exclusively Asian
haplogroup it is found nowhere else
there are no instances of it in Europe
in the new world in Africa its dominant
Tibet Japan Tibet in the Andaman Islands
very low frequencies in South Central
and Southeast Asia
it was conspicuously absent from India
it had not been seen there but recent
studies have actually found very ancient
Indian samples to be have a group D so
this is just a few of the tree after we
splitted so now this is the new D route
these are just the number of snips below
the room
this is what used to be happened room D
is M 174 and then our new branch that we
added here but with about 700 more
mutations in that block so what we
wanted to do is we wanted to have to
find more structure within this group as
opposed to just a man in his first
cousin so we took the proactive approach
and started mining our STR database and
find particular people that might belong
to this branch actually there is a well
known ft DNA sample that has had the
happy group label of de star since 2011
this came from a Nat Geo kit I think
there might have been some hesitancy
about if he was really de so the
enthusiasm wasn't necessarily there but
then we decided to take a second look
and we ran a Big Y on him and he
actually does belong to the new d2
branch but he's not closely related to
these men at all
he was aim of the 700 snips that we use
to define D - he was ancestral for over
250 of them so this is about a 25
thousand years split from a 65 thousand
year split
so the original two breakers that we
identified got a new branch called
ft 76 and so you see this structure here
we decided to do a bit more and we found
another prospective d2 member from
Hawaii again he had very limited
knowledge of his paternal history but
you can see his his autosomal results
primarily west african and he clusters
with the original two big why's that we
did but still it's a pretty ancient
relationship as he likes 58 snips than
the original to share so that we asked
me that about a 5,000 year split
so the original breakers get a new
branch so now they go from D to F T 75
to F T 76 and now they're at Ft 155 and
one last time we found another man who
had very unique STRs but didn't match
anybody closely at all he could trace
his paternal lineage to the tab
plantation so again brought over by the
slave trade is my origin showed
primarily west african and he came back
again related to this d2 lineage but he
was another twenty five thousand year
old split in this so and all we ran five
big why's we found a split at the root
of haplogroup d s maybe sixty five
thousand years ago we found three
distinct and independent lineages start
lineages twenty five thousand years old
in one of these we found another five
thousand year-old split and we
documented the first cases of haplogroup
d outside of Asia so it's actually in
the Middle East and in Africa I tried to
take a picture of what it looks like on
yet to denature e you see this giant
block this is the small block I couldn't
figure out any way to get ft seventy
five to look anywhere close to being
good enough to be shown but it'd be
about three or four times the size of
this block the information over here is
we're done up I've already mentioned
this so while we were doing this at the
same time another researcher by the name
of Hebert had actually discovered this
line as well and so he put out a paper
at the end of 2019 where he found this
new d2 lineage in Nigeria and in the
paper they referred to it as half
d0 it's similar to what was done with
haplogroup a when we found something
older than a it was referred to as a
zero found some older it's a zero zero I
saw a guy believe refers to this as as
d2 as well now the paper reference is
only three samples with a tmrc a of
about two and a half thousand years old
the the sequence data for these are
still locked up at the moment so we
don't have access to see where exactly
they fit in but there was enough
information in the supplemental material
that I was able to see where his
Nigerian samples fit in with the five
that we ran and so this is a an overview
of basically everything that we have
discovered since then this is this was
the old root of D this is the new split
so now those are d2 we have a 25
thousand year old split right here this
D to a here are the original two the man
from the tab plantation actually matches
up perfectly with these Nigerians and
this whole group shares a common
ancestor about 2500 years old but then
we have the the Russian who could traces
routes to Syria and another African
American man that are 25,000 years
removed from everybody else so there
there's still quite a bit of diversity
in this newly discovered branch and that
is my talk thanks right one point is
absolutely amazing the work that you're
doing
I'm sure there's got to be loads of
questions is this being published or how
do you communicate the readings we had
we have not published anything yet
Hebert coming out with his publication
kind of stifled that a little bit but
we're thinking of at least putting out a
note to supplement their findings with
our but we're working on that right now
of course publication takes such a long
time there's three months six months to
review and by the time the paper comes
out there's more discover all have been
made in the meantime is there any plans
for a blog or Brooke Roberta Estes wrote
a blog about it
I believe that early late 2019 when we
first added the new branch to the tree
it's something you could easily Google
Roberta Estes haplogroup D she has a
nice blog post about this so and what
extent I mean all these do this is
changing the way people think about the
evolution of men coming to what extent
do you work closely with archaeologists
with linguists who are doing the same
kind of research from an entirely
different perspective but it's obviously
very common to be very complimentary of
the work that you're doing with genetics
personally
no I'm dealing with with new data all
day every day and it's not much
collaboration outside at the moment
where would you would you like to see
that collaborations going forward or how
would how could it be coordinated it
that there could be some very
interesting things done one of the
things that we want to do is to start
analyzing more scientific samples such
as the 1,000 genomes project and other
things and we are working on that there
are there's a lot of technical
difficulties with that again it's it's a
completely different testing platform
and getting those results into our
database in a meaningful way is a bit
tricky the analysis isn't getting them
in there properly is is a bit tricky and
that was something that we discovered
discussed it in Dublin in genetic
genealogy Ireland in October and
Lara Cassidy from a Trinity College
Dublin she was presenting the results of
150 ancient genomes from Ireland they're
putting them out there really oh I mean
it's we every week it just sits there
just keep pumping them out it's it's
it's fascinating
it really is would be great to get them
into yDNA warehouse and somehow I'll
know where there are a lot of a lot of
them have a lot of them are very poor
coverage samples you can you can make a
lot of mistakes drawing inferences and
getting too aggressive with your
analysis on the ancient samples it's not
like a customer big wide where it's very
clean and and most of my decisions are
very easy you can make a lot of mistakes
yeah and and they can go unnoticed
unless the right people look at it you
have to be very careful mmm it will be
interesting to see what happens when we
finally get those ancient genomes into
the database because of course alone
will it be in radiocarbon dated and so
that will actually help us set
calibrates the mutation clock on the
tree of mankind yeah yeah exactly we
have a question here from Jaret come
thank you Mike actually I have a hundred
questions everybody allowed me one or
two so um he might watch you you're
probably familiar with heeds for
integrated mouse which what interests me
is the l21 branch when it counts for
most in ireland and we're interested in
surnames and trials and sepsis on its
world so he has taken a hundred surnames
where at least three branches contain
the same surname so we can make and
that's probably different branch for
that surname and fifty of those are
aren't related rest of Scottish Wales so
very very interesting do you have any
plans to actually do that systematically
on the big tree
know what though I'm working myself
through the limitations that privacy is
gonna is always going to be the big one
so I'm pretty sure you're referencing
alex Williamson's big tree and of course
Alex are the great tree and Contra is it
faded into the ftda the the layout or
that the formatting is is was inspired
from him just the the simple block
visualization but none of the data is
transposed or shared our trees
completely independent of the great half
would you take on the Irish flag for
instance and you get all the branches
just one thing about surnames because on
the big tree you can see the surnames of
absolutely everybody but you only see
surnames if you're within a turkey snip
distance is it possible to grow half of
surname is being put on every single
branch of the big white block tree you
know it's privacy privacy we can't
publish the names of participants adjust
the surnames I've got the person to ask
I know I'm gonna put that one thank you
very much this morning the people minute
your your backlog of
decisions and analyses of you people i
testicle
my name is freed by the hundred by the
day almost the last note the last ten
days about two thousand samples have
posted so but I'm going to routes tech
the week after so that the well you
mentioned backbone testing I haven't
talked to you about this but it may come
up with other Senate administrators I
want to buy Bob who's got a backbone
test underway you did touch on it but
could you go over the game perhaps to
explain okay I think you're dogging an
advertised thing but here we are
suddenly getting something I gather we
didn't ask for where you initiated okay
so the backbone that I'm talking about
is for people who take an STR test and
we cannot predict one haplogroup that
they are going to be in confidently so
at that time we run a few snips for them
now when it comes to the construction of
the haplit REE the the haploid tree is
automated in for the variants that were
automated automatically called by ft DNA
sometimes I make a call that ft DNA
mtDNA no calls it because they're not
confident in it but I look at it and I'm
confident it I override that and the
only mechanism I have to do that at the
moment is by uploading a results the
same way that we upload backbone results
so if you're a big wide tester and you
see a why happen the backbone order that
just means that I've been in your area
and needed to have basically uploaded a
positive snip call to get an appropriate
half a group assignment free
clear okay we'll talk with all this work
you do all of a sudden Alice's you do a
lot of it manually is automated dear own
proprietary software
what sort of applications are using to
do all this amazing work like so many
people were doing this you know it's not
so it was a lot more manual back in the
day and then some of the things have
been automated such as generating
matrixes to work with back in the day I
used to have to hand enter every single
mutation into the tree by hand enter the
snip name the position the alleles so if
I added a hundred snips and I had to do
that by hand and that would take about
30 45 minutes it was very very tedious
very recently that been able to get that
more manual then so that's just some of
the things have been automated but still
it's it's a very branches are not added
automatically everything is is viewed by
me and so every so that the analysis
portion of it is still very manual but
some of those the data collection of
business visualization from my end is a
little more automated Michael thank you
very much for the talk and I can see
every day how big the tree grows I've
even got my own particular branch me and
my dad but my question is and I also see
what really bears coming true from true
my project for all the mitochondrial
tests that are running through and I
know the fire 3 number 17 on the
mitochondrial has caused in 2016 I think
nothing's happened I was wondering a
Family Tree DNA are going to clone you
and put your other side on today
watch the mitochondria 3 thanks very
much
it honestly it's been kicked around a
bit but I don't see it happening any
time soon
it's a daunting task and it's it to be
honest I think the white tree is much
easier but no III don't see us trying to
venture down and do a similar thing with
the empty tree anytime soon my keys are
still very conservative in predicting
happy roots of most Irishmen just get or
- m269
I know for running a county project it's
very easy to see that a man is at 2 - 2
or 2 to 6 I mean you could sell about
more tests by asking average but do you
want to know I mean a descendant from
neither the by hostages are frightened
for who you are it costs about you
configure as any chance of getting more
predictions of yours so I believe it was
five six years ago FTD Nate tried to get
more aggressive with their STR
predictions and that kind of bit us in
the bud a bit because you if we get it
wrong it's on us if we predict somebody
down here and they order a test based
off of that prediction and they're wrong
well then we have to refund them comp
them and that's that's out of our pocket
so it was decided to go back we'll stick
with the basics and basically it it's
more to cover us then and the admins can
be as and and and I realize and
everybody will have their so take the
look tree and and people will say well
you can easily you can confidently
predict this and and I don't see us
going back with the more aggressive
thing any time soon either so we we just
crossed 40,000 a week ago and like I
said we've done about to that so it's
got to be around 42 now so about 42,000
well unfortunately we can stay for
another hour
we would like well unfortunately we have
to call it a day there we have Ken and
Alison Tate of talking next about
distances no object and how family
matters can be sorted through DNA but
Mike you'll be standing around the
family drinking iced down for the rest
of the day so if you have questions
please take advantage of Mike's prezi's
here and come and ask them some
questions for family treated nice done
and until then please I know you enjoyed
that