Total Pageviews

Showing posts with label SplitsTree. Show all posts
Showing posts with label SplitsTree. Show all posts

Thursday, May 23, 2013

R1a tree

Today, I want to update the STR111 tree of R1a1a that I have presented earlier here and here and here. For the first time I tried to implement some SNP information into the tree as well, which made the R1a1a branching much clearer but it is still not perfect. Assuming an initial branching of R1a 8000BP I also calculated the age of each node of the tree (see table below). Last but not least, I increased the number of individuals in this tree (N=547).

Rectangular tree of R1a (as pdf):


Polar tree of R1a (as pdf):


SNP Age in years based on tree Age in years based on STR111 variability
M420 8000 8000
SRY10831.2 7798 7907
L664 4965 4375
Z645/Z647 6117 7294
Z283 5938 6751
M458 4625 3931
L260 3598 2411
CTS11962 4013 3069
L1029 4341 3078
Z280 5614 6050
Z92 4597 3996
CTS1211 5322 5381
P278 3719 2473
CTS3402 5046 4937
L366 3079 1038
L365 4095 2041
L1280 3281 2169
Z284 5063 4688
L448 4069 2857
CTS4179 3740 2212
L176 2956 1128
Z287/Z288 4908 4499
Z93 5989 6979
Z94 5795 6900
Z2121/Z2124 5322 5319
Z2122 4124 2457
Z2123 4781 3998
L657 4729 4131
Y7 3885 2197

Due to the size of the tree I split the tree into pieces.


Edit May 25th 2013:

I did not use a molecular clock, I just defined the total age of the whole tree at 8000 years BP.

Example Z94:
1. First I calculated the standard deviation (STDEV) for each STR within Z94+ individuals (Note: For STDEV I only used individuals with STR111 data).
2. I calculated the sum of all STR111 standard deviations (STDEV DYS393 + STDEV DYS390 + STDEV DYS19+ STDEV DYS391 + STDEV DYS385a + STDEV DYS385b + STDEV DYS426 + STDEV DYS388 + STDEV DYS439 + STDEV DYS389i + STDEV DYS392 + STDEV DYS389ii + etc.). For Z94+ individuals sum of all STR111 standard deviations is 51.307. The sum of all STR111 standard deviations within R1a (all individuals of tree N=547) is 53.823.
3. I figured out a correlation between the sum of all STR111 standard deviations within Z94+ individuals and the relative age of the SNP (see formula below).
4. Arbitrarily, I defined the total age of the whole tree at 8000 years BP. It might be better to see the age estimates as relative age estimates, not as absolute age estimates.

Age of SNP Z94=8000/(2.8863*e^(0.0588*(sum of all STR111 standard deviations within all 547 R1a individuals)))*(2.8863*e^(0.0588*(sum of all STR111 standard deviations within Z94+ individuals)))

Age of SNP Z94=8000*(2.8863*e^(0.0588*51.307))/(2.8863*e^(0.0588*53.823))
=8000*(2.8863*2.71828^(0.0588*51.307))/(2.8863*2.71828^(0.0588*53.823))
=8000*59.0/68.4
=6900


Update 06/18/2013:
I generated a Z282+* STR67 tree.

Radial tree:


Rectangular tree:


Thursday, January 17, 2013

R1a1a STR111 tree

Today, I want to update the STR111 tree of R1a1a that I have presented earlier here and here. I have now collected a much larger number of samples, which makes the tree more complex. Since a lot of individuals in this tree are in the R1a1a and subclades project, I added the group name of all these individuals. Of course, this tree is based on STR data only, so the initial branching of the different SNPs can be off, e.g. Z93+ is split into multiple clusters. However, some clusters are fairly robust in terms of structure and subbranching.

Rectangular tree of R1a1a STR111 (as pdf):

Polar tree of R1a1a STR111 (as pdf):

Due to the size of the tree I split the tree into pieces. Let's look into the details:
As mentioned the initial branching in the STR111 tree is pretty messy with no clear clustering but overlapping of various SNPs. At least N86494 from Belarus with M417- ended up in the pole position. Besides very small clusters from Ireland and England and a mix of Z93+ individuals (L657+ and L657-) from Central Asia and Middle East, the first clear cluster in this tree is "4. A2. Z283+ M458+ L260+ Central European branch, West Slavic subclade", most individuals are from Poland.

Next, the narrow "9. E1. Z93+ Z94+ Z2122+ "Ashkenazi-Levite" cluster" emerges. The closest isolate individual to this R1a1a Ashkenazi-Levite cluster is the Iraqi Kurdish individual H1483. I described this close genetic proximity in some earlier posts (here and here) when STR data were still very limited, but now we have STR111 data of H1483 confirming the predicted proximity.
Other close matches to the Ashkenazi-Levite cluster are 116213 from Palestine and two Scots (162927 and 221184).

On the next two parts, we have the Z280+ individuals forming a lot of clusters, one of them is the cluster "6. J1. Z280+ CTS1211+ (CTS3402+) Southern Baltic type". Within this large mega-cluster a few Z93+ isolated individuals and isolated clusters sneaked in, e.g. the "9. C7. Z93+ Z94+ L657- Z2122-, Arabic II" cluster, and the "9. C5. Z93+ Z94+ L657-(?), Bashkirs" cluster.


Next, we have no real clustering but a mess...


Next, the nice "Z280+ CT1211+ (CTS3402-) Carpathian cluster" comes up with a clear cut between P278+ and P278- individuals.


Next, the second group of M458+ individuals appears (L260- only) forming a nice cluster. This "Z283+ M458+ L1029+ L260-" is a Central European branch.



Next, we have a mix of individuals from both major branches (Z93 and Z283) followed by CTS3402+ individuals from Eastern Europe.

Next, we have a mix, no real clustering.

Next, we have Z284+ individuals forming a cluster:
Next, we have the L664+ individuals forming a large nice cluster. Based on the current R1a SNP tree, L664+ was a very early split.  Most of these individuals are from the UK and Ireland.


Next, the Z92+ cluster emerges with people mostly from Eastern Europe.

Next, we have the Z284+ individuals forming two clusters, the Z288+/Z287+ cluster and a cluster of Z287- individuals.



Finally, the last cluster in this this tree: Z284+, L448+.


If you are getting confused by all these abbreviations, then this is okay. For clarification, I can recommend this schematic tree made by Michal.
http://eng.molgen.org/download/file.php?id=237&mode=view

Saturday, August 18, 2012

R1a1a comparison STR111 Part II

First of all, I want to thank Humata, the blogger from http://vaedhya.blogspot.com/. He helped me analyzing R1a1a; I used the same tools that he used to analyze haplogroup Q in his recent post.

So, with Humata's help I analyzed all R1a1a individuals that have STR111 data. Again, I used adjusted distances to perform this analysis, which I described and used  here and here.

In order to analyze the data I had to increase IDs to 10 digits to prevent malfunction of the used software Fitch.

There are various versions to illustrate the data. Here, I am presenting two layouts.

Rectangular tree layout (see high resolution image):


Polar tree layout (see high resolution image):




Again, to better see "what is what" I annotate each ID with the proposed group used in the R1a1a and Subclades Project at FTDNA, and I used the same colors for the subclades as in this figure from FTDNA. Even with 111 STR values the main R1a1a SNPs (Z93, Z283, etc.) are overlapping.

So what does give more accurate results, unrooted network analysis or rooted tree analysis?

It is quiet obvious that even STR111 data are not sufficient enough to differentiate between the major R1a1a subclades. There are multiple overlapping haplogroups in the tree (presented below) and in the network (presented previously), i.e. Z283 and Z93 are overlapping when focusing only on the STR111 data. All previously presented trees are adding SNP information to the tree to correct this obvious overlapping. As an example: 
179005     Krikor Mirijanian, Arapkir, Turkey who is Z93+, L342+, L657-. 

Based on his STR111 values 179005 is closest to Z283+, Z280+ and Z283+, Z284+ individuals and not to other Z93+ individuals.


Update:
After Semargl was asked how he generated his R1a1a STR111 tree and why his tree shows clear clusters along the SNP branches, he responded that he is using not only STR but also SNP information to generate the tree.  Additionally, his current phylogenetic tree is a cladogram, that means that the cladogram tree does not have any information about the age or diversity of the R1a1a branches, e.g. the Ashkenazi-Jewish Z93+, L342+, L657- cluster and the MacDonald cluster takes a large part of the tree, even though these clusters are known to be very narrow (low diversity). Hopefully, he can generate a new tree that includes all this information.



From the Fitch software manual:

In Fitch you can also randomize the input order of the sequences with option "j", jumble. Often the input order of the sequences affects the outcome of the analysis. This can be assessed by randomizing the input order. The program also asks you to specify the number of times you want to randomize the input order of the sequences. It is advisable to do jumbling at least 10 times, because it almost certainly improves the results.

This is why I repeated the analysis with 10 runs as advised. Indeed, the R1a1a STR111 tree looks a little bit better now.


Rectangular tree layout (see high resolution image):


Polar tree layout (see high resolution image):





One of the new discovered SNPs is Z1282, downstream of L342. It was found in N77532 Sundardas Tulsyan, India. He is in the tree as # "N77532-2C*". Based on the presented tree and the previously presented network analysis #184336 , SAUD ABDUL AZIZ, Qatar would be a good candidate for L1282, too (1844336-2C* in the tree and in the network analysis).
 

Monday, August 13, 2012

R1a1a comparison STR111

Today, I want to present results I got by combining two methods that I used before. The first method is described here and here; the second method is based on SplitsTree that I also used here and here.

I used all R1a1a individuals with STR111 data from the R1a1a and Subclades Project at FTDNA, a total of 203 individuals (and lots of people with unknown SNP status). The two methods don't use any SNP information, so the clusters are enterily based on STR values and the mutation rate of each STR.

Here is the network as a high resolution jpg.



Update:
To better see "what is what" I annotate each ID with the proposed group used in the R1a1a and Subclades Project at FTDNA, and I used the same colors for the subclades as in this figure from FTDNA.


Here is the color-coded network as a high resolution jpg image.


 

Tuesday, August 7, 2012

Whole Genome Comparison: Kurds vs closest genetic relatives

Similar to the previous analysis I want to using a rooted network (EqualAngle180) from SplitsTree, but in this analysis I want to include all individuals and population with an adjusted Euclidean distance of <10 towards the Kurd_D reference population of Dodecad K10a. This includes all Kurds, all Iranians and Azeris, most Georgians, a lot of Armenians, some Turks, one Iraqi Arab, and as reference populations Kurd_D, Kurds_Y, Iranian_D, Iranians (Behar), Abchazians_Y and Uzbekistan Jews: a total of 56 data points.
In this graph, the Caucasus people are closest to the root of network; KD001 and a few other Kurds are missing here (they are not part of the Dodecad project). From there the nework splits into two main branches, the left branch is dominated by Amrenians but it also includes the Iraqi Arab, Usbekistan Jews, one Georgian and one Turk and the right branch is dominated by Kurds and Iranians but it also includes two Turks. In the middle two more Turks and one Azeri show up. Interestingly, two of the Kurmanji Kurds (KD002 and KD007) are right at the root of the "Armenian" branch.

Next, I increased the number of tested samples from 56 to 112.

Again, the Caucasus people are closest to the root of network. Now, a 3rd branch in the middle appears that is dominated by Turks. The left branch now includes Assyrians and Georgian Jews.

To get a better view of the results, the same network with the Dodecad IDs.