Search This Blog

Wednesday, July 16, 2014

Model yourself as a mixture of ancient genomes


Update 12/05/2015: 4mix: four-way mixture modeling in R

...

This is really easy and should work well for most personal genomics customers (ie. those of European ancestry and with data files from 23andMe, FTDNA and AncestryDNA).

First of all, make sure you have your Eurogenes K15 ancestry proportions from GEDmatch. Then do the following:

- download the 4 Ancestors Oracle (here)

- download the Eurogenes ancient genomes datasheet (here)

- place everything into the same directory

- double click of the 4 Ancestors Oracle icon (the big red number 4)

- select the Eurogenes K15 ancient genomes datasheet

- type your Eurogenes K15 ancestry proportions into the fields provided

- hit the go button and let it rip

I'm not sure I'm allowed to upload the 4 Ancestors Oracle online, but I couldn't find the original link, so let's assume for the time being that I am. In any case, many thanks to Alexandr Burnashev for this great tool.

You'll also find some modern populations in the datasheet. They're there so that users with ancestry from outside of Europe don't end up with ridiculous results.

Obviously, you can edit the datasheet to explore more options by removing or adding individuals and populations. A spreadsheet of Eurogenes K15 population averages is available here. The oracle settings can also be tweaked in a couple of ways to fine tune the results.

If the calculator crashes, try replacing the periods with commas in both the datasheet and your ancestry proportions.

Please keep checking this post, because I'll attempt to update the datasheet at the link above every time a new ancient genome is published and has enough markers available to be tested with the Eurogenes K15. Eventually we might end up with a tool that covers most of the continents and many periods of history and prehistory.

I've done similar analyses of a variety of ancient genomes. For instance, StoraFörvar11, or SfF11, from Mesolithic Sweden came out 3/4 La Brana-1 and 1/4 MA-1, which translates to 3/4 Western European Hunter-Gatherer (WHG) and 1/4 Ancient North Eurasian (ANE), and lines up well with results reported recently for Swedish hunter-gatherers in scientific literature. You can see the full analysis StoraFörvar11 and a couple of other ancient genomes at the links below.

Analysis of Mesolithic Swedish forager StoraFörvar11

More ancient genomes from Sweden: Pitted Ware forager Ajvide58 and TRB farm girl Gokhem2

I'm still trying to answer a whole lot of e-mails so I won't be monitoring this post for a while. But please feel free to share your results and any tips you might have in the comments below.

Saturday, December 28, 2013

EEF-WHG-ANE test for Europeans


This test attempts to fit you to the three inferred prehistoric European populations as described in this recent preprint. The relevant Excel file can be downloaded here, and all you have to do is stick your Eurogenes K13 results into the fields provided to get the EEF-WHG-ANE ancestry proportions. A modified version for Near Eastern and Southeast European users can be accessed here.

The test is based on correlations between the average levels of the Eurogenes K13 and the ancient components among selected European populations. Below is a brief description of each of the ancient components.

Early European Farmer (EEF): apparently this is a hybrid component, the result of mixture between "Basal Eurasians" and a WHG-like population possibly from the Balkans. It's based on a 7500 year old Linearbandkeramik (LBK) sample from Stuttgart, Germany, but today peaks at just over 80% among Sardinians.

West European Hunter-Gatherer (WHG): this ancestral component is based on an 8,000 year old forager from the Loschbour rock shelter in Luxembourg, who belonged to Y-chromosome haplogroup I2a1b. However, today the WHG component peaks among Estonians and Lithuanians, in the East Baltic region, at almost 50%.

Ancient North Eurasian (ANE): this is the twist in the tale, a component based on a 24,000 year old Upper Paleolithic forager from South Central Siberia, belonging to Y-DNA R*, and known as Mal'ta boy or MA-1. This component was very likely present in Southern Scandinavia since at least the Mesolithic, but only seems to have reached Western Europe after the Neolithic. At some point it also spread into the Americas. In Europe today it peaks among Estonians at just over 18%, and, intriguingly, reaches a similar level among Scots. However, numbers weren't given in the paper for Finns, Russians and Mordovians, who, according to one of the maps, also carry very high ANE, but their results are confounded by more recent Siberian (ENA) admixture.

It's important to note that this test is only likely to be accurate for people of European ancestry, and indeed only those who aren't outliers from the main European clines of genetic diversity. For details of what that means, please consult the aforementioned paper. However, roughly speaking, if you're of European origin and don't score more than 3% East Asian, Siberian, Amerindian, South Asian, Oceanian, Northeast African and/or Sub-Saharan admixture, then you should get a coherent result. Users from the Near East and Caucasus should run the version specifically designed for them, while those from Southeastern Europe might find it useful to run both calculators and then compare the results.

Thanks to project member DESUK1 for putting this together at such short notice, and MfA for the modified version. Please post your results in the comments section below and state your ancestry when you do. This will help us to improve the accuracy of the test. My results make perfect sense, considering my Polish ancestry.

EEF 42.012706
WHG 40.52702615
ANE 17.46026785

This is my interpretation of who these components represent. Of course, this model might change when more ancient genomes are analyzed.

WHG and WHG/ANE: indigenous European hunter-gatherers
EEF: mixed European/Near Eastern Neolithic farmers
ANE/WHG: Proto-Indo-European invaders from the Eastern European steppe
ENA/ANE: early Uralics from the Volga-Ural region
EEF/WHG/ANE: late Indo-Europeans (ie. Celts, Germanics and Slavs)

Citation...

Iosif Lazaridis, Nick Patterson, Alissa Mittnik, et al., Ancient human genomes suggest three ancestral populations for present-day Europeans, bioRxiv, Posted December 23, 2013, doi: 10.1101/001552

See also...

Ancient human genomes suggest (more than) three ancestral populations for present-day Europeans

Ancient North Eurasian (ANE) levels across Asia

Thursday, November 21, 2013

Updated Eurogenes K13 now at GEDmatch


The new K13 population averages and genetic (Fst) distances between the inferred ancestral clusters are available here and here, respectively. To find this test at GEDmatch do this:

GEDmatch > Ad-Mix Utilities > Eurogenes > K13

Below is a 2D PCA based on the average K13 results of the European and Asian reference populations, courtesy of project member PL16.


I now have four tests at GEDmatch with Oracles: the Jtest, EUtest, K15 and K13. It's useful to keep in mind that these tests will differ in their interpretation of the data, and perhaps accuracy, depending on the ancestry of the user. For instance, the new K13 should be more useful for Central and South Asians than any of the others, because it features new reference samples relevant to them.

Monday, October 7, 2013

Eurogenes K15 now at GEDmatch


This new test is essentially an upgraded version of the EUtest. Unlike the original, it includes an Amerindian component and five native reference populations from North and Central America. So obviously it should be a lot more useful for users from the New World who are wondering about Amerindian admixture.

GEDmatch > Ad-Mix Utilities > Eurogenes > Eurogenes EUtestV2 K15

I just tried it myself, and have say that the 4-Ancestors Oracle results were impressive. In other words, they were very accurate based on what I know about my recent ancestry. On the other hand, I'd say the default Oracle was picking up more ancient gene flows. However, this might not be the case for everyone, so let's hear some feedback, discuss the outcomes, and perhaps tweak the settings if necessary.

One of the most important things to keep in mind is to ignore all results under 1%. These are likely to be noise.

The population averages and Fst distances between the ancestral clusters are here and here, respectively. Below are spatial maps of the main West Eurasian components courtesy of Gui (FR7): Baltic, North Sea, Atlantic, East Euro, West Med, East Med, West Asian.










See also...

Orcadians, the K15 and the calculator effect

Saturday, March 9, 2013

Eurogenes K36 now at GEDmatch


I've just put together a new test for GEDmatch called the Eurogenes K36. Obviously, the K36 means that it features thirty six ancestral clusters. It probably won't include any Oracles, mostly because the Calculator Effect would render these useless if they were based on the average results of the reference samples, and it'd be very time consuming for me to test a wide variety of other samples in supervised mode using thirty six sets of allele frequencies.

The main purpose of the Eurogenes K36 is to help users unravel the ethnic origins of local areas of their genomes (aka. half-segments), hence the high number of ancestral categories, some of which are very specific. In other words, the test is mainly a chromosome painting utility. It's accessible via the GEDmatch Ad-Mix link below:


GEDmatch > Ad-Mix page > Eurogenes > Eurogenes K36


An important point to keep in mind is not to take the ancestry proportions too literally. If you're, say, English, and you get an Iberian score of 12% this doesn't actually mean you have recent ancestry from Spain or Portugal. What it means is that 12% of your alleles look typical of the reference samples classified as Iberian, and this figure might only indicate recent Iberian admixture if it's clearly higher than those of other English users.

Another way to look at it is that the ancestry proportions are like map coordinates, and they'll place you with a very high degree of accuracy on a genetic map featuring other users. Indeed, please feel free to post your scores and ancestry details in the comments below to help others get an idea of what their results might represent. My results are listed below. The scores put me squarely in Poland relative to those of other European samples I've run, which is correct.


Also worth mentioning is that this test focuses on much deeper ancestry than the Ancestry Composition at 23andMe. Hence, I expect that many Europeans will score a few percent in non-European clusters. However, like many ADMIXTURE results, this could give us strong hints about population movements into Europe during prehistory and early history, so it's worth keeping an eye on.


Monday, December 3, 2012

4-Ancestors Oracle at GEDmatch


The Jtest and EUtest at GEDmatch now include a new tool called the 4-Ancestors Oracle (aka. Oracle-4), as well as the 3D PCAs I promised earlier. Oracle-4 will attempt to pinpoint your ethnic group of origin, and then also work out the most likely combinations of two, three and four ancestral populations which make up your genome. However, this doesn't mean the results will actually show your ethnic group, or those of your parents (in dual mode) or grandparents (4-way mode). They might for many people, but for others they'll reflect the best possible outcomes from the reference samples available.

GEDmatch Ad-Mix Utilities

Enjoy, and feel free to give feedback to John at GEDmatch if you think it might be useful (but please don't spam his account).


Thursday, September 27, 2012

Jtest K14 - the Eurogenes Ashkenazi ancestry test


Update 19/03/2018: It's come to my attention that many people are still using the Jtest and taking the results very seriously. Indeed, perhaps too seriously.

Also, some users are doing weird stuff with the Jtest output in an attempt to estimate their supposedly "true" Ashkenazi ancestry proportions, like multiplying their Ashkenazi coefficient by three, because Ashkenazi Jews "only" score around 30% Ashkenazi in this test. Ouch! Please don't do that!

Let me reiterate that this test was only supposed to be a fun experiment. It was never meant to be the definitive online Ashkenazi ancestry test. And even as fun experiments with ADMIXTURE go, it's now horribly outdated, and probably useless for anyone with less than 15-20% Ashkenazi ancestry.

So it might be time to move on. If you really want to confirm your Jewish ancestry, either or both Ashkenazi and Sephardi, then you need to look at much more powerful and sophisticated options. One of these options is the Global25 analysis (see HERE), which can pick up minor Jewish ancestry of just a few per cent. But it's not free (USD $12), and it's a DIY test that requires a bit of time and effort to get the most out of it. Also, you'd need to send me your autosomal file so that I can estimate your Global25 coordinates. But I can help you get started and even quickly check if you have any hope at all of confirming Jewish ancestry.

If, for whatever reason, you'd rather not take advantage of the Global25 offer, because, say, you don't want to share your data with me, then it might be an idea to join the Anthrogenica discussion board and ask the experienced members there about other options [LINK].

In any case, whatever you choose to do, please remember the following points, and feel free to share them with others who are still using the Jtest:

- do not multiply your Jtest Ashkenazi score by 3 in an attempt to find your "true" Ashkenazi ancestry proportion, because this won't work for the vast majority of users

- but do compare your Jtest Ashkenazi score to those of other people of the same or very similar ancestry to yours to get a rough idea whether you might have any Ashkenazi ancestry (the Jtest population averages will be useful for this, see here)

- if you're still not sure what your Jtest results mean, then just focus on your Jtest Oracle-4 output at GEDmatch, and if you don't see AJ at the top of the oracle list, then this is a strong signal that you don't have substantial Ashkenazi ancestry
...

I recently learned that the new Ancestry Painting at 23andMe will include an Ashkenazi reference group. To be honest, I’m not sure there’s much value in using a genetically bottlenecked population of varied biogeographical origins as a reference in such things. Indeed, the Ashkenazi mainly descend from a few hundred founders, but carry Central European, Eastern European, Middle Eastern, African and probably many other admixtures, as evidenced by their genome-wide and uniparental markers.

That’s quite a problem, because due to their relative inbreeding, they produce strong ancestral clusters in many analyses, like in ADMIXTURE runs. However, these clusters are made up of allele frequencies from a wide range of sources and, paradoxically, it’s the relatively more outbred populations which contributed to the Ashkenazi gene pool at its formative stages that often end up showing Ashkenazi admixture in such tests, despite not having any. I've seen this happen regularly in my experiments with ADMIXTURE and STRUCTURE, and I'm pretty sure I could find an example in a peer reviewed study if I tried.

That’s just how things work with the algorithms we have available to run these sorts of tests. Nevertheless, since 23andMe is incorporating an Ashkenazi cluster into its new painting, I thought I’d try and come up with an Ashkenazi ancestry test to perhaps get a rough idea of what we might expect. I'm using ADMIXTURE in supervised mode, and basically trying to recreate clusters that have shown up in a variety of fine-scale analyses, including my ChromoPainter run of Northern European samples. It’s still a work in progress, but below are links to files that many of you might find useful..

Jtest K14 files

Jtest averages for selected populations

EUtest K13 files

EUtest averages for selected populations

The Jtest folder contains files that can be used to make an Ashkenazi ancestry test/chromosome painting with 14 Eurasian and African clusters. The EUtest folder contains the same files, except that the Ashkenazi allele frequencies have been removed. It’s useful to cross check results from both tests, mainly to see what’s hiding under the Ashkenazi admixture if it shows up in the Jtest.

Based on a few test runs today, I’d say that the noise level for the continental clusters is much less than 1%. But it rises to a few per cent for the intra-West Eurasian clusters. In other words, if you’re European, then you might score something like 0.02% in the Sub-Saharan cluster, which basically means 0%. However, you might get around 2% in the Middle Eastern cluster, even though you’re from Central Europe, and you don’t have any recent Middle Eastern ancestry. You can blame various prehistoric and historic migrations into Europe for these seemingly quirky results, and also the fact that Mesolithic Europeans were significantly Eurasian (i.e. Siberian, Amerindian and South Asian-like).

The Ashkenazi cluster is very similar to the Middle Eastern cluster in that regard. So anyone who gets an Ashkenazi score of around 2-3% either has very distant Jewish ancestry or, more likely, none at all. However, those who show more than 25% membership in that cluster are almost certainly of fully Ashkenazi ancestry, and their genomes peppered with Ashkenazi-specific chromosomal segments.

There’s really not much difference between 2% and 25%, you might say. In fact, there is if we say there is. As always, the main thing to remember is that these clusters don’t really exist, because genetic variation is clinal, so the cluster names are basically arbitrary and it’s always the relative results that matter.That’s why to really understand what your scores mean, you need to compare them with those of other users.

Obviously, it's best to compare with people from the same ethnic and/or regional groups. If the Ashkenazi + East Med scores look relatively inflated, that's a sign of recent Ashkenazi ancestry.


Feel free to use the files above for anything you want, except commercial stuff. Please note, I make no guarantees that they’ll provide accurate results for everyone. I might update this post early next week with new and/or additional files and more tips.

...

Update 6/10/2012: The Jtest K14 and EUtest K13 will soon be available at GEDmatch, accompanied by an "Oracle" population matching analysis and maybe even a 3D genetic map. If all goes to plan, the population matching test should be able to give a decisive yay or nay to anyone wondering whether they have recent Ashkenazi ancestry.

By the way, below is a PCA based on the Jtest averages for selected populations. It was produced by one of my project members so that we could check the reliability of the 14 "ancestral" components. The samples were classified into clusters based on their highest peaking component. So, for instance, the Scots are in the light blue Atlantic cluster, along with French Basques, because the Atlantic component dominates in both groups. However, overall, they're more similar to other samples than to each other.

As per above, the plan is that GEDmatch will soon offer a 3D genetic map based on the loadings from this PCA analysis.


Update 11/10/2012: The Jtest and EUtest are now on offer at GEDmatch. The quickest way to get there is via this link to the Ad-Mix page. Then, from the drop down menus, choose Eurogenes, followed by Jtest.

First run the Admix test to check whether your Ashkenazi admixture is significantly higher than expected for your part of the world (as per above, Jtest averages for selected populations are available here). Then move on to the Oracle analysis by pressing the relevant button at the bottom of the page.

If your Ashkenazi admixture is clearly elevated, and the top 20 single and/or mixed mode Oracle results show AJ (Ashkenazi Jews) as one of your potential matches, then it’s likely you have recent Ashkenazi ancestry.

Whether that’s the case or not, you can then move on to the Chromosome Painting feature to see where the potential Ashkenazi admixture is located in your genome. It’s useful to cross check the results with those from the Ancestry Finder at 23andMe to assess their accuracy.


As already mentioned, the EUtest is exactly the same as the Jtest, but with the Ashkenazi allele frequencies taken out. You can use this option to see what’s hiding under your Ashkenazi admixture in the Jtest. To compare your results with those of selected populations from Europe, Asia and Africa, refer to the EUtest averages sheet.

Please note: it's important to interpret the results with insight. You need to learn how the system works, pay attention to the types of populations that appear in your results, consider carefully why they might be paired with other populations, and of course study the statistics in detail. Expecting a bullseye classification at the top of the Oracle list is likely to lead to major disappointment for many people, simply because I don't have enough samples to represent all of the substructures that exist around the world, especially within countries.

I’ll try and update both tests in a few weeks, after seeing how successful the whole set up is at predicting Ashkenazi admixture and locating it in the genome. One of the main goals will be to improve the accuracy of the Oracle analysis for everyone, including New World people with Amerindian admixture.

Update 21/10/2012: Below are spatial maps of a few of the ancestral clusters from the Jtest, courtesy of project member FR7.


Update 4/12/2012: The Jtest and EUtest at GEDmatch now include a new tool called the 4-Ancestors Oracle (aka. Oracle-4), as well as the 3D PCAs I promised earlier. Oracle-4 will attempt to pinpoint your ethnic group of origin, and then also work out the most likely combinations of two, three and four ancestral populations which make up your genome. However, this doesn't mean the results will actually show your ethnic group, or those of your parents (in dual mode) or grandparents (4-way mode). They might for many people, but for others they'll reflect the best possible outcomes from the reference samples available.

Enjoy, and feel free to give feedback to John at GEDmatch if you think it might be useful (but please don't spam his account).