Total Pageviews

Sunday, July 8, 2012

Harappa Ancestry Project presents Kurdish data

Zack from the Harappa Ancestry Project was so kind to focus on the Kurdish samples in his monthly update of the project and compared them with available autosomal Kurdish data.

I used his results to calculate the Euclidean distances of the different Iranian groups in his data set.
Based on his analysis, the Kurds in the Harappa Ancestry Project are closest to Iranians.

Here are the TOP30 matches for Kurds in the Harappa Ancestry Project:


The biggest surprise to me is one Bulgarian participant (HRP0209) at #17 in the Top30 list of Harappa Kurds. I am not sure how to interpret his small distance to other Kurds/Iranians.

Top30 matches for HRP0209 are:

Update: Mystery solved. 
HRP0209 is not Bulgarian, but Kurmanji Kurd from Turkey (KD006). For some reasons, he prefers to be mislabeled in various projects (e.g. "BGTR1" on Eurogenes).

Saturday, July 7, 2012

R1a1a L657+

In my previous post I used my approach to look into Y-STR67 values of individuals of the FTDNA R1a1a and Subclades Y-DNA Project.

Today, I want to risk a deeper look into L657+. (2. D. Z93+ L342+ L657+ Central&Southwest Asia). Here are the Top30 matches of the L657+ Modal haplotype:

 
Obviously, the L657+ cluster is not narrow at all. The closest L657+ individuals to the  L657+ Modal haplotype are from the Arabian peninsula (one exception: N2358 Philip, Ayroor, Kerala, India).

L657+ shows high variance, a lot of L657+ are not in the Top30 matches of the modal haplotype. For some of those we have additional information about the location:

M6986 (#52),
M7443 (#61),
N102178 (#74),     Pakistan
N22414 (#80),       India
M6736 (#81),        Iran
M7417 (#82),
U2810 (#83),         Pakistan
112208 (#94),        Kazakhstan
M6740 (#95),        Saudi-Arabia
216619 (#104),      Pakistan
N12617 (#119),     India
U2321 (#128)        India

I want to discuss the bolded ones a bit.

N102178    (Mehrdad, Lahore/Pakistan).
N102178 is clear outlier.


N22414    (Luddan Singh Ranu, 1800s, Manki, Punjab/India)
N22414 shows some similarities to U2321 who is also from Punjab/India.

 M6736    (Iran)
M6736 is similar to L657+ individuals from the Arabian peninsula (one exception: N2358 Philip, Ayroor, Kerala, India).

 U2810    (Tharn Bajwa 1890 Pakistan)
U2810 shows small similarities to L657+ individuals from the Arabian peninsula
112208    (Babasan tribe from Northern Kazakhstan)
112208 is a clear outlier.


M6740    (Mecca/Saudi Arabia)
M6740 shows small similarities to L657+ individuals from the Arabian peninsula.

 216619    (K Khan Qureshi, c.1830-1900 Pakistan)
216619 is a clear outlier. Interestingly, he shows some similarities to one individual, 209438 from Bitlis close to Lake Van. Unfortunately 209438 did not test for L657.


N12617    (Jagarnath Dixit ca 1490-1550 India)
N12617 is a clear outlier.


U2321  (Amar Sandhu, Jalandhar, Punjab, India)
U2321 shows some similarities to N22414 who is also from Punjab/India (see above).

Conclusions:
1. The L657+ group does not show a clear narrow cluster when looking at Y-STR67 values.
2. Just the individuals from the Arabian peninsula show a cluster.
3. The L657+ individuals of the "Al Tamimi" and the "Al Rass" family (Saudi-Arabia) seem to be very closely related.
4. The high number of "Al Rass"/"Al Tamimi"  data in the FTDNA L657+ group (5/22) is shifting the modal haplotype towards Saudi Arabia. Thus, I am proposing an adjusted modal haplotype for L657+.

Changes are listed below:



DYS391 DYS389i DYS389ii Y-GATA-H4 DYS607 CDYa CDYb DYS534 DYS444
FTDNA 10 13 30 11 17 36 40 13 14
Adjusted 11 14 31 12 16 35 41 14 13


 Here are the Top30 matches of the adjusted L657+ Modal haplotype:

Now, the closest to modal haplotype is the L657+ individual N2358 Philip, Ayroor, Kerala, India followed by L657+ individuals from the Arabian peninsula ("Al Rass"/"Al Tamimi" families).


Still, a lot of L657+ are not in the Top30 matches of the modal haplotype but they moved up in the ranking (new number in italic):

M6986 (#52==>#39),
M7443 (#61==>#5),
N102178 (#74==>#54),     Pakistan
N22414 (#80==>#9),       India
M6736 (#81==>#37),        Iran
M7417 (#82==>#59),
U2810 (#83==>#105),         Pakistan
112208 (#94==>#46),        Kazakhstan
M6740 (#95==>#62),        Saudi-Arabia
216619 (#104==>#110),      Pakistan
N12617 (#119==>#123),     India
U2321 (#128==>#67)        India


Friday, July 6, 2012

How to read STR data

I know that there is a lot of controversy about the usage of Y-STR data and I agree with most of them. However, sometimes there are no other data available (e.g. Y-SNP data) and in those cases Y-STR can help a little bit to understand the observed pattern within one haplogroup subbranch.
Looking at STR databases (e.g. ysearch.org or semargl.me/en/dna/ydna/tools/asd-classic/)  can be painful and useless, the reason for this is that the relationship between two individuals is solely based on STR differences, but these differences are not "weighted" in any sense, they just focus on "Distance markers" and "Distance steps".

Everyone who took a look at one of the FTDNA project quickly realizes that some Y-STRs are more variable than others In the L342+ group of FTDNA R1a1a and Subclades Y-DNA Project the following order of variance can be observed in the first 25 Y-STRs (from low to high variability):

DYS388    DYS437    DYS392    DYS455    DYS448    DYS393    DYS454    DYS426    DYS447    DYS438    DYS390    YCAIIb    DYS459a    DYS385a    DYS389ii    CDYa    DYS389i    YCAIIa    DYS464c    DYS449    DYS19    DYS464d    DYS464a    DYS385b    DYS459b    CDYb    DYS460    Y-GATA-H4    DYS391    DYS607    DYS439    DYS456    DYS570    DYS458    DYS576    DYS442    DYS464b

Differences in DYS464b are more common than differences in DYS388, so differences in DYS464b are "less important" than differences in DYS388. Any ranking should be weighted according to the variability of the STRs. The variance of some Y-STRs was calculated and published previously; YHRD listed them here.

This additional information can be used to better rank the best matches for an individual. The less Y-STRs are available for a comparison the more this approach is useful.

Of course, I did a first ranking test using a Kurdish individual:
H1483 (Z93+, L342+, L657-) in the FTDNA R1a1a and Subclades Y-DNA Project (focusing on 34 STRs and individuals that have these 34 STRs tested).

Top30 matches:



Next, this approach was expanded using 67 Y-STRs of the Arabic modal haplotype (2. C6. Z93+ L342+ L657-, Arabic). Top30 matches are:


Obviously, the Arabic cluster is a very narrow one with low variance. It cannot be old.


Then, R1a1a Ashkenazi-Levite modal haplotype was tested:


Similar to the Arabic cluster the Ashkenazi-Levite cluster is also pretty narrow with low variance. It cannot be old, either.



Wednesday, July 4, 2012

Genetic impact of the biggest cultural invention (agriculture)

In a previous post I presented how domestication of animals impacted small segments of the human genome very specifically. Today, I want to present how it changed the overall human genepool of Eurasia and beyond. From genetic studies we now know that not only the domesticated animals but also the farmers spread from the Middle East to the surrounding areas. But why is it that in some regions, a different paternal haplogroup and a different autosomal component is dominating?

Here, I am presenting two models, one for the paternal haplogroups, and one model for the autosomal components based on Dodecad K12b. Please, don't try to read distances or geographic directions out of the model, they depict the idea in an rough schematic manner.

Paternal haplogroups:


Dodecad K12b components:

Sunday, July 1, 2012

Ashkenazi-Levite Jews and their Iranian origin

Ashkan
This is an Iranian name that is still in existence, for thousands of years. And I believe the term "Ashkenazi" is ultimately derived from the Iranian name "Ashkan".


More than 2000 years ago after the eastern conquests and death of Alexander the Great, the Parthian Empire emerged stretching from Balochistan to the Levant.
Photo: Courtesy of Encyclopaedia Britannica
 
The Parthian Empire is also known as the 'Arsacid Empire' named after the first Parthian king Arsaces I of Parthia. In Iranian languages (Farsi/Kurdish), his name was Ashk (اشک) or Ashkan (اشکان) and his Parthian Empire is still called 'Ashkanian' (اشکانیان) in Iranian languages.

According to the bible a 'Kingdom of Ashkenaz' existed. Together with Ararat, Minni (Mannaeans), and the Medes they prepared for war against Babylon (Jeremiah 51:27-28). It is possible that 'Ashkenazi' became a generic term for Jews who migrated northward into the Parthian Empire, given the assumed location of the Kingdom of Ashkenaz.
The proposed lineage tree of the Ashkenaz is also described in the bible: Ashkenaz belongs to the "Japhetic" branch, not to the Semitic one. BTW, the Hebrew language belongs to the Semitic language family.

In the bible Noah had three sons: Shem (Semitic), Ham (Hamitic), and Japheth (the rest of the known world). So, the word Shem refers to both, Noah's son and the father of all Semitic people. Gen.10 declares the following nations as semitic.

"Gen10:22 The children of Shem; Elam, and Asshur, and Arphaxad, and Lud, and Aram. "

Asshur and Aram refer to Assyrians and Arameans, respectively.

Noah's other son Japheth had also children.
Gen10:
2 The sons of Japheth; Gomer, and Magog, and Madai, and Javan, and Tubal, and Meshech, and Tiras.
3 And the sons of Gomer; Ashkenaz, and Riphath, and Togarmah."

So Ashkenaz is Gomer's son, and Gomer is Japheth's son.

Gomer is regarded as a synonym for people who lived in Anatolia in that time, either Cimmerians, Gauls in Anatolia (Galatians/Celts) or people from Cappadocia, the literature has not decided yet.
Ashkenaz is also regarded as synonym for Skythians or Saka. Much much later in Medieval Times, 'Ashkenaz' was the Hebrew word for 'Germany'.

The ancestors of Ashkenazim may actually lived in Northern Mesopotamia for many centuries before moving into Europe, especially into the Rhineland as merchants. Given that most Ashkenazi Jews left Israel after the Roman conquest and they do not seem to appear in Europe until Charlemagne's time (800 AD), they must have lived somewhere else prior to that. I think it is reasonable that the ancestors of Ashkenazim were Aramaic speakers who lived in the Parthian Empire.

So, is it possible to see some Iranian admixture in Ashkenazi Jews of today?

Let's take a look at the 2. C2 Ashkenazi-Levite (Type "A" / "AJ") modal haplotype, R1a1a Z93+ L342+ L657-:

I will take a first look using 37 STRs:
DYS393    DYS390    DYS19    DYS391    DYS385a    DYS385b    DYS426    DYS388    DYS439    DYS389i    DYS392    DYS389ii    DYS458    DYS459a    DYS459b    DYS455    DYS454    DYS447    DYS437    DYS448    DYS449    DYS464a    DYS464b    DYS464c    DYS464d    DYS460    Y-GATA-H4    YCAIIa    YCAIIb    DYS456    DYS 607    DYS 576    DYS 570    CDY a    CDY b    DYS 442    DYS 438

Here are the Top30 matches for this R1a1a Ashkenazi-Levite modal haplotype (excluding all actual Ashkenazi-Levite Jews).


Within the Top30 you can find plenty Iranians including the Iraqi Kurdish individual H1483 that was also described here.


Conclusion:
The Ashkenazi-Levite modal haplotype for R1a1a Z93+ L342+ L657- may be Iranian in origin.

Thursday, June 28, 2012

Kurdish Y-DNA Part VII

Now, we have another two individuals with haplogroup J2 (highlighted in yellow):

Information about Kurdish autosomal DNA has been updated:
HarappaWorld
Dodecad K12b
Eurogenes K12b
McDonald

1x E1b1b1c1a (Alevi Kurmanji from Dersim/Turkey)
1x G2a (Alevi Kurmanji from Turkey)
2x J1 (Feyli, originally from Iran)
1x J1c3 (Sorani from Iran)
1x J2 (Zaza from Dersim/Turkey)
1x J2 (Kurmanji from Dohuk)
1x J2 (Kurmanji from Turkey)
1x J2a3a (J2a1a at 23andme; J2a4a at ISOGG 2009; he is M47+, M322+)(Yezidi from Iraq)
1x T (Sorani from Koysinjaq/Iraq)
1x R2a (Sorani from Sulaymaniyah/Iraq)
1x R1b1a2* (Kurmanji from Zakho/Iraq)
1x R1b1b2a (Zaza from Turkey)
1x R1b1 (P25+)(Kurmanji from Maras/Elbistan/Turkey)
1x R1a1a (Z93+, L342+, L657-)(Sorani from Sulaymaniyah/Iraq)
1x R1a1a (Z283+, Z282+, Z284-, M458-, Z280-, subclade 3  only his paternal great-grandfather is Kurdish from Turkey)
1x R1a1a (Alevi Zaza from Dersim/Turkey)
1x R1a1a (Alevi Kurmanji from Dersim/Turkey)
1x I2a2a* (old I2b1*; L1229-, L1230-, L1226-, L699-, L701-, L702-, L703-, L704-, M379) (Sorani from Sulaymaniyah/Iraq)

So far all tested SNPs of the I2a2a* individual turned out be negative.

More data can be found here:
Kurdish Y-DNA Part I
Kurdish Y-DNA Part II
Kurdish Y-DNA Part III
Kurdish Y-DNA Part IV
Kurdish Y-DNA Part V
Kurdish Y-DNA Part VI

Kurdish Y-DNA Part VIII

mtDNA of Kurds V

 Just an update (new entries are highlighted in yellow):

1x C4b (Alevi Kurmanji)
1x G2a (Sorani)
1x H5a1 (Sorani)
1x H13a2 (Alevi Kurmanji from Dersim)
1x H14 (Yezidi)
1x H15a1 (Sorani; mtDNA fully sequenced here and here)
1x H15b (Sorani)
1x HV (Sorani)
1x HV (Kurmanji from Zakho)
1x I5a (Zaza from Dersim)
1x J1b (Sorani)
1x J1c (Alevi Kurmanji from Dersim)
1x J2a1a  (Kurd from Turkey)
1x N1b1 (Alevi Kurmanji from Dersim)
1x U1a1 (Zaza)
1x U1a1 (Sorani)
1x U5a1 (Kurmanji from Dohuk)
1x U8b (Feyli)

Some information about the latest entries:

mtDNA G2a:
The mtDNA haplogroup G2a is not known in populations from the Middle East but in Ainu from Japan and in Northeastern Siberia close to Bering Strait. There is only one reported case of a fully sequenced mtDNA G2a from the Caucasus (Georgian). However, this Georgian individual is G2a1b (not the same subbranch). Another curiosity is that this Kurdish individual has the 16172C mutation. 16172C was only described once in a Chinese individual. Strange enough, the Chinese individual does not belong to G2a but to the neighbor haplogroup G2b (to be precise G2b1a). My guess is that
a) it is a very, very rare coincidence to have the same mutation as two indepedent events,
b) 23andme is not testing for G2b mutations, or
c) the mtDNA G haplogroup tree needs to be updated.
Behar et al. 2012 estimated an age of 26788 ± 4618 years for G2.
Behar et al. 2012 estimated an age of 17146 ± 5270 years for G2a.
Here is a map of 23andme users with mtDNA G2a created by Evon_Evon who has a general interest in this mtDNA haplogroup.

mtDNA I5a:
The mtDNA Haplogroup I is found in Europe, Middle East and South Asia. Quintana-Murci et al., 2004, analyzed the mtDNA of 20 Kurds from Iran, and found one individual with the mtDNA I (1/20=5.0%). Several cases of I5a (Germany, Italy, Romania, Russia, USA) are described in the FTDNA mtDNA Haplogroup I Project.
Behar et al. 2012 estimated an age of 15116 ± 4128 years for I5a. Fully sequenced I5a mtDNA is available from Yemen, Dubai, and Turkey (GenBank). The individual from Turkey is the closest match for the Zaza (both are lacking the "Arabian mutations "G3705A","T5096C", and "G5773A"). 

mtDNA U5a1:
The mtDNA Haplogroup U5a1 is quiet common in Europe but it is also present in other parts of Eurasia being one of the most common and best described mtDNA haplogroups. Full mtDNA sepuences of the root of U5a1 are available from England, Spain, Caucasus, and Czech Republic at GenBank. The mtDNA of famous Cheddar Man from England (who lived 9,000 years ago) turned out be U5.