Total Pageviews

Friday, August 24, 2012

Haplogroup J1 tree STR111

Edit 09/12/2012:
I updated this post to make it shorter/easier for readers.

Today, I want to present the haplogroup J1 tree with STR111 data. I used the same method as before. Most of the individuals in this tree are from the Arabian peninsula and have the haplogroup J1b2b1 (aka J1c3d2) L222.2+. Some of them did not test for L222.2 but I am pretty confident about this. Based on the tree I predicted the haplogroups of some others.

Unfortunately, there is no Kurdish data included in this tree but we can assume that most Kurds with J1 belong to the same subclade as their closest neighbors, i.e. J1* M267. Of course, this has to be confirmed.

Actually, one individual in this figure is definitely from Iraqi Kurdistan: Irq---92829 is an Assyrian of Erbil :-)

I know of one Kurd of Sharif descent that is J1b2b. Unfortunately, he has only STR67 analyzed, so I cannot include him in the STR111 trees presented below. However, he is grouped with 100329 (unknown origin) and M4284 (from UAE) in the FTDNA J Haplogroup Project, so he would be close to these in the tree, and both 100329 and M4284 are at the root of the very large J1b2b branch (The previous name of J1b2b was J1c3d; at 23andme it is called J1e; details about the name changes can be found at ISOGG).

Color code:
J1 M267 Z1834-, Z1842- grey ("oldest" branch)
J1 M267 Z1842+ light blue
J1a M365.1 red (only one individual: Antonio Gomes 1635, Milhazes, Barcelos, Portugal [Por---73612]) 
J1b* L136  light green
J1b2* P58 dark green (clover)
J1b2b* L147.1 blue
J1b2b* Jewish Cohanim Cluster orange
J1b2b*  L859+ yellow
J1b2b1 L222.2+ purple    


Here are the results.
Rectangular tree of J1 (as jpg and as pdf):


Polar tree of J1 (as jpg and as pdf):

Based on this STR111-based tree the history of J1 is much clearer. 

1. The ancestral region of J1 is Caucasus, Northern Mesopotamia or Eastern Anatolia.
2. There is an early split of J1 into two branches.


Edit 09/24/12:
Rob1 (eng.molgen.org user) helped me collecting more J1 STR111 data. The new tree has a total of 313 users (see below). Three Jewish clusters emerged in the STR tree, all three clusters are highlighted in light orange, orange, and dark orange. Arab clusters within the blue J1b2b* L147.1 area are highlighted in light brown, brown, dark brown, and asparagus green. The greenish Arab cluster is at the root of the J1b2b1 L222.2+ purple cluster. The most obvious cluster within the J1b2b1 L222.2+ purple cluster is Bany Zaid cluster (highlighted in red).

Edit 09/30/12:
Iyyovi (eng.molgen.org user) asked me to upload pdf versions of the latest J1 rectangular and polar tree.
pdf polar tree
pdf rectangular tree

Thursday, August 23, 2012

Haplogroup J2 tree STR111

Today, I want to present the haplogroup J2 tree with STR111 data. I used the same method as before (The annotation is not ready...)

Here are the results.

1. Rectangular tree of haplogroup J2:

2. Polar tree of haplogroup J2:
Edit:
A better J2 tree.

Tuesday, August 21, 2012

Whole Genome Comparison: Kurds Part II

 I recently saw the ACD tool (ACD=Ancestral Component Dissection) at Vaêdhya. Unfortunately, I was not able to download it, but I could see the logic behind it by looking at the presented figures. What he did was:
1. Taking several populations from one region and assuming only one main ancestry for all of them.
2. Taking the lowest percentage of each component (Dodecad, Harappa or Eurogenes) and declaring this lowest percentage as ancestral.
3. Subtracting the actual percentage of a component of a population from the lowest percentage.

Example with the Baloch component in Harappa Project:
Step 1:
Armenian: 18%
Assyrian:  19%
Kurd: 26 %
Iranian: 27 %
Step 2:
Lowest Percentage for the Baloch component in these four populations: 18%
Step 3:
Armenian: 18% - 18% = 0 %
Assyrian:  19% - 18% = 1 %
Kurd: 26% - 18% = 8 %
Iranian: 27% - 18% = 9 %

His goal is to measure the influx migration from outside. This measurement is based on the assumption of an ancestral population for all compared populations. Let's see how this ancestral population would look like:

The minimum values of the 12 components for Armenians, Assyrians, Kurds and Iranians are:
S Indian         0%
Baloch         18%
Caucasian    42%
NE Euro        1%
SE Asian       0%
Siberian         0%
NE Asian       0%
Papuan           0%
American       0%
Beringian       0%
Medit.            5%
SW Asian    11%
San                0%
E African      0%
Pygmy          0%
W African    0%
Total           77%

Let's bring up the numbers to a total of 100% :
S Indian        0%
Baloch        23%
Caucasian   55%
NE Euro       1%
SE Asian      0%
Siberian        0%
NE Asian      0%
Papuan          0%
American      0%
Beringian      0%
Medit.           6%
SW Asian    14%
San                0%
E African      0%
Pygmy          0%
W African    0%
Total         100%

In short:
#PopulationPercent
1 Caucasian 55
2 Baloch 23
3 SW-Asian 14
4 Mediterranean 6
5 NE-Euro 1

The main drawbacks of this approach are that ...
1. The calculation is based on processing of already processed data, so the error is getting bigger. It is not directly based on raw data.
2. It is based on the assumption that neighboring people with different histories/languages/religions have one main common ancestry.


I actually used a similar approach here but I did not use averages of populations and I did not assume a common ancestry for these populations. Instead I only focused on results of individuals of only one population. My goal was not to see differences but to see similarities. Onur critized that my old approach was not based on admixture results of others and not on raw data.


So, what I now did was to take all Kurdish raw data and fuse them into one genome. This genome is based on the allele frequencies in Kurds results that I presented here. Allele frequencies of 25-75% were considered as heterozygous and allele frequencies below 25% and above 75% were considered as homozygous.

Then, I used this fused Kurdish genome and ran it through HarappaWorld, Dodecad K12b, and Eurogenes K12.

Interestingly, the Harappa results of the fused Kurdish genome are not far away from the results above:

HarappaWorld for fused Kurdish genome (N=16):

#PopulationPercent
1 Caucasian 54.63
2 Baloch 28.84
3 SW-Asian 12.61
4 Mediterranean 2.85
5 NE-Euro 1.07


Closest populations to this fused Kurdish genome (N=16):

Single Population Sharing:


#Population (source)Distance
1 kurd (yunusbayev) 9.97
2 kurd (xing) 10.36
3 assyrian (harappa) 11.41
4 armenian (yunusbayev) 11.69
5 kurd (harappa) 12.08
6 azerbaijan-jew (behar) 12.29
7 armenian (behar) 12.83
8 iranian (harappa) 12.93
9 uzbekistan-jew (behar) 13.06
10 georgian (harappa) 13.52
11 iranian-jew (behar) 13.91
12 iraqi-mandaean (harappa) 14
13 georgia-jew (behar) 14.49
14 azeri (harappa) 14.62
15 iranian (behar) 15.3
16 turkish (harappa) 16.27
17 iraq-jew (behar) 16.47
18 turk (behar) 17.64
19 turk-kayseri (hodoglugil) 18.51
20 kumyk (yunusbayev) 19.35


Dodecad K12b for fused Kurdish genome (N=16):

#PopulationPercent
1 Caucasus 52.25
2 Gedrosia 29.22
3 Southwest_Asian 12
4 Atlantic_Med 4.18
5 North_European 2.35

Single Population Sharing:


#Population (source)Distance
1 Kurds (Yunusbayev) 10.3
2 Kurd (Dodecad) 11.48
3 Armenians_15 (Yunusbayev) 11.7
4 Iranian (Dodecad) 11.95
5 Azerbaijan_Jews (Behar) 12.09
6 Uzbekistan_Jews (Behar) 12.59
7 Assyrian (Dodecad) 13.02
8 Armenian (Dodecad) 13.25
9 Georgia_Jews (Behar) 13.38
10 Iranian_Jews (Behar) 14.49
11 Iranians (Behar) 14.87
12 Armenians (Behar) 15.19
13 Turks (Behar) 16.88
14 Iraq_Jews (Behar) 17.74
15 Turkish (Dodecad) 19.37
16 Kumyks (Yunusbayev) 20.68
17 Druze (HGDP) 23.08
18 Adygei (HGDP) 23.34
19 Lezgins (Behar) 23.49
20 Chechens (Yunusbayev) 23.62


Eurogenes K12b for fused Kurdish genome (N=16):


#PopulationPercent
1 Caucasus 45.67
2 W-Central Asian 22.79
3 Mediterranean 19.29
4 Southwest Asian 11.81
5 North European 0.45




Sunday, August 19, 2012

I2a2a*-M233 (old I2b1*) comparison STR111

Today, I want the present the calculated tree of I2a2a*-M233 (old I2b1*) using the same approach as before, 10 randomized runs to improve the tree.
Quiet frankly, the result of the M233 STR111 analysis is disillusioning and disappointing. None of the subbranches shows clustering in the tree, all subbranches are overlapping. So based on STR111 values nothing (at least within M233) can be predicted, maybe in the future with more M233 individuals having STR111 tested.

Anyways, I want to share the current results:



Saturday, August 18, 2012

R1a1a comparison STR111 Part II

First of all, I want to thank Humata, the blogger from http://vaedhya.blogspot.com/. He helped me analyzing R1a1a; I used the same tools that he used to analyze haplogroup Q in his recent post.

So, with Humata's help I analyzed all R1a1a individuals that have STR111 data. Again, I used adjusted distances to perform this analysis, which I described and used  here and here.

In order to analyze the data I had to increase IDs to 10 digits to prevent malfunction of the used software Fitch.

There are various versions to illustrate the data. Here, I am presenting two layouts.

Rectangular tree layout (see high resolution image):


Polar tree layout (see high resolution image):




Again, to better see "what is what" I annotate each ID with the proposed group used in the R1a1a and Subclades Project at FTDNA, and I used the same colors for the subclades as in this figure from FTDNA. Even with 111 STR values the main R1a1a SNPs (Z93, Z283, etc.) are overlapping.

So what does give more accurate results, unrooted network analysis or rooted tree analysis?

It is quiet obvious that even STR111 data are not sufficient enough to differentiate between the major R1a1a subclades. There are multiple overlapping haplogroups in the tree (presented below) and in the network (presented previously), i.e. Z283 and Z93 are overlapping when focusing only on the STR111 data. All previously presented trees are adding SNP information to the tree to correct this obvious overlapping. As an example: 
179005     Krikor Mirijanian, Arapkir, Turkey who is Z93+, L342+, L657-. 

Based on his STR111 values 179005 is closest to Z283+, Z280+ and Z283+, Z284+ individuals and not to other Z93+ individuals.


Update:
After Semargl was asked how he generated his R1a1a STR111 tree and why his tree shows clear clusters along the SNP branches, he responded that he is using not only STR but also SNP information to generate the tree.  Additionally, his current phylogenetic tree is a cladogram, that means that the cladogram tree does not have any information about the age or diversity of the R1a1a branches, e.g. the Ashkenazi-Jewish Z93+, L342+, L657- cluster and the MacDonald cluster takes a large part of the tree, even though these clusters are known to be very narrow (low diversity). Hopefully, he can generate a new tree that includes all this information.



From the Fitch software manual:

In Fitch you can also randomize the input order of the sequences with option "j", jumble. Often the input order of the sequences affects the outcome of the analysis. This can be assessed by randomizing the input order. The program also asks you to specify the number of times you want to randomize the input order of the sequences. It is advisable to do jumbling at least 10 times, because it almost certainly improves the results.

This is why I repeated the analysis with 10 runs as advised. Indeed, the R1a1a STR111 tree looks a little bit better now.


Rectangular tree layout (see high resolution image):


Polar tree layout (see high resolution image):





One of the new discovered SNPs is Z1282, downstream of L342. It was found in N77532 Sundardas Tulsyan, India. He is in the tree as # "N77532-2C*". Based on the presented tree and the previously presented network analysis #184336 , SAUD ABDUL AZIZ, Qatar would be a good candidate for L1282, too (1844336-2C* in the tree and in the network analysis).
 

Monday, August 13, 2012

R1a1a comparison STR111

Today, I want to present results I got by combining two methods that I used before. The first method is described here and here; the second method is based on SplitsTree that I also used here and here.

I used all R1a1a individuals with STR111 data from the R1a1a and Subclades Project at FTDNA, a total of 203 individuals (and lots of people with unknown SNP status). The two methods don't use any SNP information, so the clusters are enterily based on STR values and the mutation rate of each STR.

Here is the network as a high resolution jpg.



Update:
To better see "what is what" I annotate each ID with the proposed group used in the R1a1a and Subclades Project at FTDNA, and I used the same colors for the subclades as in this figure from FTDNA.


Here is the color-coded network as a high resolution jpg image.


 

Sunday, August 12, 2012

L342+ comparison STR67

Today, I want to present results I got by combining two methods that I used before. The first method is described here and here; the second method is based on SplitsTree that I also used here and here.

I used all L342+ individuals with at least STR67 data from the R1a1a and Subclades Project at FTDNA, so L657- and L657+ individuals (and lots of people with unknown SNP status).

Here is the network as a jpg. (Update: New results of N101746 are included in jpg link.)
I started to annotate the different clusters and I will finish it soon, still I want to share the current status.

Update:
N101746 (Central&Southwest Asian; India) is in a cluster with:
M6851 (Arabic II; Saudi-Arabia),
M6699 (Arabic II; Unknown),
M6698 (Arabic II; Unknown),
M6853 (Arabic II; Unknown),
197670 (Central&Southwest Asian; India),
M6132 (Central&Southwest Asian; UAE), and
160543 (Central&Southwest Asian; Iraq).
The "Arabic II" individuals are very close to each other, while the "Central&Southwest Asian" individuals in this cluster show more diversity.

This described cluster above is close to another cluster that is very narrow. The follwoing individuals belong to this cluster:
157103 (Arabic; Saudi-Arabia),
160271 (Arabic; Qatar),
162855 (Arabic; Saudi-Arabia),
178907 (Arabic Saudi-Arabia),
178905 (Arabic; Saudi-Arabia),
157621 (Arabic; Saudi-Arabia),
157621 (Arabic; Saudi-Arabia),
157619 (Arabic; Saudi-Arabia),
M6895 (Arabic; Saudi-Arabia),
178906 (Arabic; Saudi-Arabia),
M7066 (Arabic; Unknown),
M6183 (Arabic; Unknown),
M6679 (Arabic; Unknown),
M6458 (Arabic; Kuwait),
M7013 (Arabic; Kuwait),
M6982 (Arabic; Unknown), and
M6285 (Arabic; Qatar).

Zoom in: