Friday, 22 April 2016

A Comparison of Big Y Analysis Resources

There are several very useful resources available to us for interpreting the results of our Big Y tests. Here is a brief summary of what we get and what we don't from each resource - all have their Pros and Cons and all add something to the overall interpretation of the results. Because this is a completely new area of science, and we are on the crest of the wave of scientific discovery, the different analyses from the different sources frequently produce different results, which in turn allow us to ask why and refine our methodologies further. This will continuously change over time as we understand more, adjust our SNP declaring criteria, and refine our interpretation of the data. We can expect big changes to take place over the next few years.


FTDNA
The presentation of the Big Y results from FTDNA is currently quite limited. This is not surprising as they were the pioneers in this area and their first offering in terms of how the data is presented is now outdated. But this is due to change in the near future when they introduce their new Big Y features. What these are as yet we don't know but we can expect exciting developments over the next few months. The march of progress carries on! 

Currently we are given a list of close matches and the number and nature of Shared Novel Variants with each match (i.e. shared new SNPs), also Known SNPs that are not shared, and SNPs that are unique when comparing just two specific individuals. FTDNA also places us on their own version of the Haplotree (the human evolutionary tree) which tells us what SNPs lie at branching points above our own particular sub-branch.

Confusion arises from a number of different issues, some of them general points, some of them FTDNA-specific:
  • the separation of new ("Novel") SNPs from those already identified ("Known")
  • FTDNA's high threshold criteria for declaring a SNP misses some SNPs
  • often no SNP names are reported, only SNP positions - you have to go to YBrowse.org or YFULL to get specific information regarding the name of SNPs present at a particular location on the Y chromosome.
However, and most importantly, FTDNA provide the facility to download our raw data (in .vcf, .bed, and .BAM files) which allows us to have the data analysed and interpreted by a host of other resources. Details of how to access and share these files can be found here. The "Download VCF" option will download the .vcf and .bed files (about 1.6 mB) and the "Share BAM" option will allow you to copy a temporary link to your BAM file (which is >600 mB in size and so is far too big to be sent by email).

My current position on FTDNA's haplotree with details of SNPs tested
(green, positive; red, negative)
(click to enlarge)


YFULL
YFULL gives us a more detailed analysis of BAM files and places us on their Y-Haplotree in relation to other people nearby (i.e. who have undergone NGS [Next Generation Sequencing] testing, like the Big Y). It also identifies our terminal SNP (or SNP block), the SNPs at branching points further upstream (and hence the Shared SNPs we have with our neighbours), and the unique / personal / private SNPs that each member possesses (currently).

Over and above FTDNA's analysis, YFULL tells us the following:
  • how many people on adjacent branches have tested and where they are from
  • SNP names and any "equivalent SNPs" (i.e. exactly the same position on the Y chromosome but alternative name)
  • time estimates for the formation of each SNP (and hence the particular branching point)
  • TMRCA estimates for the people in each sub-branch (with 95% Confidence Intervals)
  • information on each SNP including position on the Y, ancestral and derived values, alternative names, and reference sequence. This information can be supplemented by YBrowse.org
  • easier-to-understand presentation of currently unique (personal) SNPs with an estimate of their "quality"
  • data on about 500 STRs including the majority found in FTDNA's 111 marker panel - this can be helpful in calculating TMRCA estimates and narrows the 95% range around the estimate (compared to TMRCAs based on 111 marker data)

The Gleeson Lineage II portion of the YFULL haplotree
(click to enlarge)


Full Genomes (FGC) Analysis
The FGC analysis of BAM files is comparable to the YFULL analysis but their public Y-Haplotree is not as user-friendly as others and is of limited utility. FGC generates the following reports with the underlying files (processed BAM file, mtDNA, and STRs):
  • A detailed analysis of called variants report
  • A variant genotyping report
  • Haplogroup classification
  • Y-STR report, and 
  • mtDNA report

clarifYdna analysis
clarifYDNA will reanalyse your Big Y data for $30 and produces a Y-DNA haplotree report from your results - this will be periodically updated as new data becomes available from other testers.  Reports are in sync with a recent version of the ISOGG haplotree, and are able to indicate which aspects of the phylogenetic structure are robust and which are more tenuous - thus it combines aspects of the Y-haplotree that are both "established" and "provisional / experimental". 

Unfortunately the tree is only available to subscribers and is not available to the public.

Click here for an Example analysis.


Haplogroup Project Administrators 
The Administrators at the Z255 Haplogroup Project, and indeed the Admins of more upstream Haplogroup Projects (e.g. L21, R1b & subclades, etc) are an incredible resource and their respective Yahoo Discussion Groups are great places to post questions and get replies.

John Murphy puts together a regular updated spreadsheet / haplotree for the Z255 group which has an advantage over YFULL's analysis - it incorporates new SNP discoveries from specific SNP Packs and single SNP testing and not just from NGS testing (Big Y, FGC).

The Gleeson Lineage II portion of John Murphy's spreadsheet
(click to enlarge)


Alex Williamson's "Big Tree"
Alex is one of the most important people in the R1b research community and is a champion of data analysis, interpretation, and (most importantly) presentation. His Big Tree website (www.ytree.net) is a masterful display of complex data in a digestible format. He places us on his haplotree so that we can see our terminal SNP (block), SNPs at upstream branching points, and our neighbours on adjacent branches.

Advantages over YFULL include:
  • The Big Tree gives us our neighbours names and places of origin, thus making it easier to form an impression of where a particular sub-branch might have formed and if it is specific for a particular surname.
  • Easy to navigate with lots of additional information by simply clicking on a surname or a SNP.
  • His graphics are superb.
  • His Mutation Matrix allows us to see which SNPs are shared and which SNPs are not between us and our closest neighbours. 
  • His presentation of unique (personal) SNPs gives us not only an estimate of "quality" but also the region of the Y-chromosome in which they are found (this can be useful in judging if this is a true SNP and also how easy it would be test for it in a bespoke SNP Panel)

The (current) 4 branches of the Gleeson Lineage II on the Big Tree


Nigel McCarthy's Z255 Subgroup
Nigel is another pioneer. He is one of the first people to combine SNP markers and STR markers into a single tree. We are lucky enough in the Gleeson Lineage II group that we are closely related to  some of the people in Nigel's McCarthy DNA Project. As a result, Nigel has included us in the Z255 portion of his phylogenic tree (Group E).

Nigel's own SNP analysis is complementary to the ones above and he too will occasionally discover new SNPs that others have not included in their analyses.

Major advantages over previous analyses include:
  • As well as the SNPs, he also presents STR data and the change in STR values at each branching point
  • He includes people who have not been tested on the Big Y (i.e. anyone with Y-DNA-67 or Y-DNA-111 results). As a consequence, his portion of the haplotree contains more Gleeson's from Lineage II than any other tree - 12 members altogether (compared to 9 members on Alex's tree, 9 on John Murphy's, and 6 on the YFULL tree).
  • From Nigel's analysis it is possible to see where Back Mutations and Parallel Mutations occur in the STR markers.


The Gleeson Lineage II members in Nigel McCarthy's Group E of his McCarthy DNA Project


Mike W’s Haplotype Data for R1b-L21
Mike is an administrator of several FTDNA projects and a leader in the genetic genealogy community for a long time. He maintains a very comprehensive spreadsheet that can be downloaded from the Links section of the R1b-L21 project Yahoo group or a smaller version from the Z255 Yahoo group. This spreadsheet collects the STRs for 67 and 111 markers, and SNPs from the Big Ys or other sources. A user can calculate his genetic distance in relation to the complete database and he can infer his haplotype according to the most common haplotype of his closest matches. The spreadsheet also calculates the group mode and several statistics required to characterize a particular group.


James Kane's SNP Matrix
I am new to James' SNP Matrix but it too is a work of art, a magnum opus, not surprising for a 90 mB spreadsheet. Yet again, James' approach to NGS data analysis offers a fresh perspective and can detect possible/probable SNPs that have not turned up in other analyses. Having multiple analyses and interpretations of the same data is a great advantage - it allows us to see points of agreement and points of difference in the various approaches, and ultimately helps us to question the data more intelligently which in turn will lead to better analysis and interpretation.

James’ matrix compares the SNPs of all the participants while the other methods preselect the relevant SNPs. This capability is very important for the identification of new potential SNPs. These can be checked against other analyses for consistency or disagreement. Additionally, it is possible to evaluate unique SNPs or SNPs that belong to a particular group with different levels of quality or if they are part of the combBED area. It also provides positions for both the build 37 (GRCh37) and 38 (GRCh38) human reference genome sequence.

The matrix workbook requires BAMs for inclusion. What the scripts do is visit each file for every variant location and outputs the read depth in a very large combined VCF file. The idea is to remove the ambiguity of BED files. The old HTML based pages did include everything, but became unwieldy. They are being replaced soon.

James also has a blog site and an Experimental Y Tree (currently being updated) with SNP names & their equivalents, surnames, places or origin, and TMRCA estimates.

Everyone who has done the Big Y test should send James a link to their BAM file so he can include you in his analysis. This looks at the data from yet another perspective and helps further with the interpretation. It should help clarify the discoveries from other sources and may even identify some additional SNPs.

Below are the instructions and an explanation that James has put together about sending him your BAM file for analysis and what will happen after that:

What’s needed?
A link to your Big Y BAM file using the “Share BAM” button on your Raw Results page.  Let me know if detailed instructions would be helpful.  Please also include this statement in the email:

As the owner or administrator of FTDNA kit#, [YOUR KIT#], I consent to allow analysis of the Y DNA contained in the provided BAM file.  The results of this analysis may be used the phylogenetic tree of haplogroup R or independent researchers for scientific purposes.

What will be done?
Your BAM will be downloaded and realigned to GRCh38.  This will allow a new VCF/BED to be created and compared with others.  Results will be included in http://www.it2kane.org/matrix/R-P312.html.  When sufficient analysis is available for the branches, it will be possible to include time to most recent common ancestor estimation based on these results.  The new VCF/BED will be provided to those interested.

What won’t be done?
Unlike the commercial 3rd party analysis, you won’t get mtDNA, STR value estimates, or variant naming any time in the near future.

See below for my data use policy.

Raw Data Policies

In light of the recent dust-up between FTDNA and another 3rd party site, I have codified my data usage policies.
1.     VCF/BED files submitted for analysis are made available for other R-L21 researchers using the R1b-L21(S145) Haplogroup and Subclades Y DNA forum hosted on Yahoo.  This aids researchers to correctly assign variants to their related haplogroups.
2.     Raw BAM data is retained in a password protected cloud storage account.  The project recognizes there is a low probability that files may contain data not actually on the Y chromosome, which may reveal medically relevant information about the participant. BAM files may be individually shared with qualified researchers and analysts only after approval of the sample’s owner.
3.     GRCh38 aligned versions of variant calls and BED coverage generated by the project’s bio-informatics workflow can be shared with researchers without the sample owner’s explicit authorization.
4.     FTDNA kit #’s are displayed for convenience of related surname projects or haplogroups in all reporting.  As the identifier is used to log into the FTDNA account this has security implications for the kit owner.  Project members may request reporting on tree or call matrix reports use an internal project id instead.
5.     Project members have the right to request that their raw data is removed from reporting at any time, but shared variants in the tree will be retained.


Some Closing Thoughts
This is a new science and we are still trying to get to grips with it. The pithy saying "Many hands make light work" operates quite nicely in this situation. It is only by looking at the data from a variety of different perspectives that we can hope to understand it better, and quickly. So we should be using all of the above utilities to analyse and interpret our Big Y results. Thanks to the internet, this process (which previously would have taken decades to complete) can now be accomplished in a matter of years thanks to what effectively is a crowd-sourcing approach - a group of citizen scientists working together toward a common goal and employing the power of the internet to communicate and collaborate effectively.

There is still a lot of testing to be done - we need more people to do the NGS tests (Big Y, FGC tests, etc). And we need clever people to develop more tools for analysis, interpretation, & presentation of the data. But as this critical mass of people tested builds, and as our ability to analyse and interpret  and present the data improves, we will begin to reap greater and greater dividends. 

Software packages are being developed to help build combination family trees using SNP data, STR data, and standard genealogy. Already you are able to add DNA markers to your Family Tree on Ancestry. This will advance even further and trees will start to be linked online via downstream SNP markers.

Furthermore, for Irish genealogies at least, we will be able to link some of our family trees to the Ancient Annals and Genealogies, bringing us back to before the time of surnames, back to 900 AD, 800 AD, 700 AD, and even further.

In a few years, when we look back at this time in human history, we will be able to say ...
I was there. 
I contributed to that. 
I helped build the Evolutionary Tree of Mankind. 
And I know exactly where I sit on it.


Maurice Gleeson
German Creamer
Lisa Little
April 2016







Monday, 18 April 2016

Lineage I News - Big Y Results

Thanks to Alex Williamson for his work on the world tree, and thanks to our three members who are pioneers on the frontier of genetics by taking the Big Y test to find their place on Alex's Big Tree. Through their efforts, the Gleason Lineage 1 group is now known to be located far downstream of SNP DF27 in our own unique block (R-6650231-C-T). Finding this characteristic SNP block is analogous to finding our stem (or twig) on the genetic tree. Those closest to us, on nearby stems of the branch, have roots in England, Scotland, and Ireland.

Lineage I members on Alex Williamson's Big Tree

The 7-digit number identifying the block indicates the position on the Y-chromosome where there is an anomaly (mutation) not found among others in the general population. The C-T designation says that the base chemical "C", which is usually found in that location, has morphed into a "T" in our family. You might say that this block of mutations, which are unique to us, define our temporary characteristic or terminal SNP. But as more people test, it may not be unique to our family, and we may need to look further downstream to locate our uniqueness. 

Another fascinating discovery is that two of the three subjects share an additional SNP downstream from the block discussed above, at position 22117264-C-T. This can be seen on that portion of the Big Tree that Lineage 1 occupies: http://www.ytree.net/DisplayTree.php?blockID=1368&star=false. By moving right to left and clicking one of the “parent” SNPs at the top of this page, you can expand the tree to see more and more upstream branches. 

Current SNP progression for Lineage I

Keep in mind that each member has mutation locations that are his alone (so far). It is the shared mutations that are shown in the blocks, and these are not found (as yet) in others. If more members of our lineage were to take the Big Y test, it might be possible to identify the sub-lineages among the family by their shared SNPs.

Judith Gleason Claassen
April 2016





Thursday, 3 March 2016

MDKA Profile – Patrick Treacy 1794-1878 (G-80, B38804)

Name of MDKA: Patrick Treacy

Family nickname/agnomen: not known

Date & location of birth: circa 1794, probably Carhoon, Tynagh, Galway

Date & location of death: 14 Nov 1878, Sheeaunrush, Portumna, Galway

Residence: Carhoon, Tynagh, Galway, Derrew, Kilquain, Galway & Sheeaunrush, Portumna, Galway

Occupation: Farmer

Other family occupations: ...

Religion: Roman Catholic

Wife's name: Mary Egan

Date & location of wife's birth: circa 1811, possibly Kilemore, Killimor, Galway

Date & location of wife's death: 25Oct1877, Sheeaunrush, Portumna, Galway

When & where married: circa 1828, Galway

Sponsors at wedding: unknown

Order of children's birth:

1 - Edward circa 1831
2 - Patrick 07Jan1833, Killimor
3 - Mary 12Mar1835, Killimor
4 - John 20Jun1836, Killimor
5 - James 30Sep1839, Killimor
6 - Nicholas 13Mar1843, Killimor
7 - Mary Anne 24Feb1846, Killimor
8 - Bridget 07Jan1849, Killimor
9 - Michael 25Mar1852, Killimor

Possible name of father & mother (based on naming convention): Edward & Mary Anne

Sponsors at children' baptisms (numbered, and in order of children's birth):

1 - unknown
2 - Michael Gorman & Mary Egan
3 - Thomas Costelloe & Mary Treacy
4 - Martin Keane & Catherine Eagan
5 - Ferdinand Ridge & Anne Egan
6 - Edward Trasy & Honora Egan
7 - Ferdinand & Mrs. Ridge
8 - Patrick Treacy & Mary Cain
9 - Patrick Trasy & Catherine Martin

DNA kit number(s): YSEQ: 1870; FTDNA: B38804

Project ID number(s): G-80

DNA tests taken: 
23&Me - autosomal
YSEQ - R1b-L21 Super-Clade Orientation Panel, R1b-Z255 Panel, A558
FTDNA - Autosomal DNA Transfer; Y-DNA37; A557; Z255; Big Y; Y-DNA67; Y-DNA111

Terminal Y-SNP sequence: Z255 > Z16439/7 > A557 / A558 > Z29008

Closest match(es) ID No (with Genetic Distance in brackets): At 67 Markers - 60393 (6); N101540 (7); 244645 (7); 338070 (7); 395582 (7)





Friday, 19 February 2016

The Truth about William and Abiah Gleason

(or Why I Learned to Look for the Original)

William Gleason was the fourth son of Thomas and Susanna (Page) Gleson. Thomas and Susanna were natives of Suffolk, England, who first appeared at Watertown in the Massachusetts Colony in 1652. William’s birth year can be approximated as 1648 based on his appearance as a witness in court records of 1671 when he is said to be 23 years of age.[i] It is unknown whether he was born in England or in the colonies. If his birth were in Watertown, there would be no baptismal record since the Watertown church records prior to 1686 are lost. For the same reason, there is likely no church record for the birth of William’s wife, or of their marriage, if these events occurred in Watertown. Although his parents moved from Watertown to Cambridge, then to Charlestown, and back to Cambridge, William lived in Watertown with his uncle William Page throughout his youth. For his services to the Page family he received a legacy of £10 in the will of William Page, to be paid at age twenty-two.[ii]

The name Abiah or Abia is a biblical name, as is Abiel, and Abijah; and the name of the wife of William can be found spelled as each of these variants. However Abiel is always a masculine name in the Bible, but Abiah and Abijah appear as both feminine and masculine. Since the name Abiah was often used for girls during that period in the colonies, Abiah is the spelling assumed here.

The Record Book of the Pastors of Watertown Church is available beginning in 1686. On 10 April 1687, an entry to these records tells of the baptism of four children of Abiah Leason [sic] who held the covenant. She is described as the “…wife of young William Leason.” The names of the children are William, Joseph, John, and Elizabeth.[iii] On 31 July, Abiah Leason [sic] was admitted to “Full Communion” in the Watertown Church. From these records, we have the knowledge that the wife of William was named Abiah, was a member of the church, and that she was the mother of his children. A birth record for only one of these children can be found. Because the couple settled at Cambridge Farms (now a part of Lexington), the birth of the oldest son of William and Abijah Gleison [sic] is entered on Cambridge records as 15 April 1679.[iv] Two more children, Esther and Isaac, were born to the couple after 1687.[v]

The facts are straightforward to this point, but now the genealogists enter the picture, confusing the facts with misinformation and transcription error.

The errors in the White genealogy
were propagated for decades

The compiler of the first genealogy of the Gleason lineage was John Barber White, a genealogist who never revealed a source, and published his genealogy in 1909.[vi] Many have quoted it as near Biblical truth. Of course we are grateful for his admirable efforts in searching vast records in a day when that was a very difficult thing to do. He created a guidebook of possibilities that we can use as a starting point for our research. Unfortunately, he employed the “best guess” technique to arrive at some conclusions. This method was used to determine William’s birth place and year. Moreover, when he found a girl named Abiah Bartlett, born 1651 in Watertown, she became the best guess for the wife of William.[vii] And it came to pass that Abiah Bartlett was accepted as his wife, and this fiction has been published willy-nilly on the Internet in countless Gleason family pedigrees. Had White looked more thoroughly into early records he would have discovered that Abiah Bartlett actually married Jonathan Sanders/Sanderson in Cambridge in 1669.[viii]

Note: Besides the aforementioned errors in William’s birth data and the error in the identity of his spouse, White is guilty of another major error concerning this family: William had no daughter Ann as he claims. According to the original church record, the Ann Leason [sic] in question was admitted as a young adult to the Watertown Church on 22 Jan 1687/8, where she held the covenant as an unmarried woman who lived with her mother.[ix] This Ann was actually the youngest child of Thomas and Susanna (Page) Gleason.


A confusing misprint crept into Torrey's book

The next genealogist to enter our discussion is Clarence A. Torrey, a man highly respected in the field and renown for his work: New England Marriages Prior to 1700.[x] He began his compilation in the 1920s and continued for years using only paper and pen. In 1985 the hand-written work was first transcribed and put into a printed version. Since that time the New England Historic Genealogical Society (NEHGS) has published the version used today, and a search on their website for the marriage yields this later printed version. Unfortunately the NEHGS entry for William and Abiah has incorrectly placed a right parenthesis, leaving one to wonder if Abiah Gleason might have later married Sanderson.[xi]


To resolve the question, the only thing to do was to seek out the original manuscript of Torrey. When I requested a look at the original, NEHGS graciously provided a scan of Torrey’s written notes. He obviously was aware of the error in White’s Genealogy, and he intended to correct White by making a notation that William’s wife was not Bartlett [double underlined in original]. He further wrote that “she married Jonathan Sanderson,” in reference to Bartlett. He gave William’s wife’s name as “Abiah ____” which is explained, in the introduction to the book, to mean that her surname is unknown. The introduction also explains that when the date of a marriage is unknown, the date of the birth of the first child is used, preceded with “by” meaning before. Thus he gives “by 1679” as the date of marriage, which in no way implies that the marriage year was 1679. Torrey’s original got it right, and the printed result should be:



The moral of the story is that conflicts in genealogy can sometimes be resolved with a look at the original. Of course the original is not always available, and we must rely on material that is someone’s interpretation of the handwritten. So be aware when quoting any genealogical source­ - the truth may yet remain hidden.


Torrey's original handwritten notes reveal the truth (last entry)
(click to enlarge)

Judith Gleason Claassen
Feb 2016





[i] Middlesex County, Massachusetts, Abstracts of Court Files 1649-1675, (Online database. AmericanAncestors.org. NEHGS), 2:137.
[ii] Middlesex County Probate, 16345.
[iii] Massachusetts Vital Records to 1850, Watertown: Record Book of the Pastors, 120, 121-122. (Online Database, AmericanAncestors.org, NEHGS). The name Gleason is often spelled as Leason throughout early church records, but civil records use a variation of Gleason. If doing a search, one should try both spellings.
[iv] Massachusetts Vital Records to 1850, Vital Records of Cambridge.
[v] Ibid.; Record Book of the Pastors, 129.
[vi] John Barber White, Genealogy of the Descendants of Thomas Gleason of Watertown, Massachusetts, 1607-1909 (Haverhill, Mass: Nichols Print, 1909, repr. Bowie, Md., Heritage Books, 1992).
[vii] Ibid., 29; Massachusetts Vital Records to 1850, Watertown.
[viii] Massachusetts Vital Records to 1850, Cambridge
[ix] Record Book of the Pastors, 126. Listed as Ann Leason.
[x] Torrey’s New England Marriages Prior to 1700. (Online database.  AmericanAncestors.org. NEHGS) Originally published as: New England Marriages Prior to 1700. Boston, Mass.: New England Historic Genealogical Society, 2015.
[xi] Ibid.












Wednesday, 17 February 2016

Lineage II MHT - John Murphy's version

Apart from Alex Williamson's Big Tree version of the Mutation History Tree (MHT) for Lineage II (see previous post), there are several other versions and I'll cover each of these over the next several weeks. 

It is very important, and very helpful, that several different people are each offering their own interpretation of the same data and generating their own particular version of the tree based on their interpretation. One can compare the different versions and identify any differences in interpretation which in turn leads to discussion and an exchange of scientific rationales. Ultimately, this will lead either to a consensus opinion or to a qualified disagreement, with each party knowing and understanding the rationale behind the other party's position. So far, there have been no major differences in the various versions of the Lineage II MHT generated by the different people involved.

John Murphy's Z255 Mutation History Tree

The first alternative version is that by John Murphy. John is one of the Co-Administrators for the Z255 Haplogroup Project and all members of Lineage II are encouraged to join this particular project. The project has 566 members and is for anyone who has tested positive for the SNP Z255, as have the members of Lineage II. Below is a reminder of the abbreviated SNP progression for Lineage II:
R … > M269 > L150 > L23 > L51 > L151 > P311 > P312 > L21 > DF13 > ZZ10 > Z255 > Z16437/9 > Z16438 > BY2852/3/4 > A5629

Neal Downing is the group Administrator with John Murphy, Kim Fields and Mike Walsh as Co-Administrators.  They also run a very informative Yahoo Groups Mailing List (the Z255 and Subclades Forum) and again all Lineage II members are encouraged to join this as there are some very talented people involved in discussions on this list.

Below is the relevant portion of John's version of the tree. Lineage II members can be found in the lower left-hand side.

(click to enlarge)

The format of John's tree is different to Alex's version but it is based on the same data, with one important difference - Alex only includes NGS test results in his tree (i.e. Big Y, FGC tests, 1000 Genomes Project) but John also includes people who have undertaken single SNP tests and SNP Packs such as the Z255 SNP Pack, which was introduced by FTDNA late in 2015. This is most evident on the far right of the tree (orange box).

As a result, John's version of the tree contains more individuals than Alex's version.

Both versions only show the Shared SNPs that two or more people in the overall group have in common with each other - Unique SNPs are not shown in either version of the tree.


Maurice Gleeson
Feb 2016





Monday, 15 February 2016

Big Y SNP Markers of Lineage II

In the previous post, we reported on the three new sets of Big Y results, where these new results placed those members in Alex Williamson's Big Tree  and how these new data changed the overall structure of our own little Lineage II portion of the human evolutionary tree. I include the diagram again below, just to recap.

In this post, we will take a closer look at the actual SNP markers  - both the Shared SNPs within the group, and the Unique SNPs for each individual.

On the left, the 9 Lineage II members who tested on Big Y & their relation to each other

Shared SNPs among Lineage II members

The diagram above contains only the SNPs that are shared between individual members - therefore, I call this the "Shared SNP" portion of Alex's tree. But in addition, each member has unique SNPs, particular to that specific individual, that are not shared with anyone else in the world (at least for the time being ... but that is liable to change and I will explain why below). We'll take a look at the Shared SNPs first and the Unique SNPs afterwards.

Alex produces fabulous tables and matrices of the results on his Big Tree website and these are really informative.  The table below shows all 9 members of Lineage II who have undertaken the Big Y test. The positions of the relevant "Gleeson" SNPs on the Y chromosome are shown in the first column, and the SNP name (if they have one) is in the second column. Many SNPs will not have a name because they are only newly discovered and no one has got around to naming them yet! This is "cutting-edge science" after all - we are on the crest of the wave of new scientific discoveries and there will not be immediate answers to all our questions. Of the 26 SNPs in the table, only 7 have been named (currently).

The third column tells us what "block" the SNP belongs in, followed by the specific region of the Y chromosome where it is found. And the rest of the columns are the actual data for our 9 members, with the kit number, family name, & terminal SNP (or terminal SNP block) for each member at the top of each column.

It is important to appreciate what a mammoth task this is. The table below represents a distillation of approximately 26,000 SNPs from each of the 9 members of Lineage II. The SNPs in the table below are only shared among these 9 members in Lineage II and by no other of the 2000+ people who have undergone NGS (Next Generation Sequencing) testing with the Big Y and similar tests. [1]

To see the table above, go to this link (www.ytree.net/DisplayTree.php?blockID=16), find the Gleeson portion in the bottom right of the diagram, click on the A5629 SNP block, and on the next page click on Show Mutation Matrix.

Taking the first row of data as an example, all 9 members are positive for this SNP, which lies at position 18606400. Furthermore, the usual "base" found at this position in most people is C (cytosine) but in our case it is a T (thymine). The SNP name is A5629 and it sits within the A5629 SNP block (each block is usually named after the SNP that appears at the top of the list of SNPs within it, which in this instance is A5629 - you can see this in the tree diagram at the top of the page).

There are two different types of classification system in operation in these tables. Some cells have a white, pink or grey background; and on top of this they may contain +, *, **, or ***. Here's an explanation for these classification systems:
  1. The colour of the background gives an indication of the degree of "coverage" achieved by the particular Big Y test. In other words, the Big Y test measures the entire Y chromosome in small chunks, usually 30-100 times per chunk, but (purely by chance) some chunks are read only a few times and (purely by chance) some chunks are not read at all. So a white background indicates good enough coverage, a grey background indicates no coverage, and a pink background indicates questionable coverage. These pink regions often indicate that the individual may be positive for a SNP even if it does not show up in his data. This is important to appreciate because it has a very significant implication: just because someone does not test positive for a given SNP doesn't mean it is not there! 
  2. The +, *, **, and *** signs mean slightly different things for the Big Y test and the FGC tests. In very crude terms, I like to think of them as "definite, probable, possible, and unlikely". In other words, they give some indication of the probability that the SNP in question is a genuine SNP (and not a "false positive"). But that's just a rough guide. You can read Alex's more comprehensive explanation below. [2]
Applying these classification systems to the SNP markers in the table, you can see that 15 of them are "definite" (a + on a white background) and the rest of them are unlikely (***). If you wanted to put a probability on how likely is "unlikely" ... if you said 10% you might not be far off. In other words, "unlikely" SNPs may be genuine SNPs 10% of the time, and "false positives" 90% of the time. Hopefully time (and further testing) will tell. But for now, this raises a second important point worth appreciating: just because someone tests "positive" for a SNP does not mean it is there!

By comparing which SNPs are shared among which members, it is possible to separate out members who are more closely related to each other than to other members in the group. And in this way it is possible to construct a family tree for these people based on their SNP markers. This produces a Mutation History Tree based on SNPs alone, which is essentially what Alex's tree is - a SNP-based Mutation History Tree.

Two things will happen over time: 1) some SNPs will be reclassified (in terms of "definite, probable, possible, and unlikely"), and 2) as more people test on the Big Y (or similar tests), some of the SNP blocks will be split into smaller blocks or individual SNPs, just like we saw when member 411177 (Glisson) joined the project - the addition of his results made SNP A5628 split away from the rest of the A5629 block (see the previous post). 

In other words, as more people test, the tree will branch and subdivide. And because we currently have 16 SNPs in 4 SNP blocks, it is likely we will have an additional 16 branches added to the tree.


Unique SNPs among Lineage II members

If we go to the Lineage II portion of Alex's tree (on the right side of the diagram here) and click on any of the individual names, this launches a table with that particular individuals unique SNPs. The same classification systems for SNPs apply in terms of coverage (white, grey, pink) and probability of being genuine (+, *, **, and ***).

From the Tables below for each of the 9 individual members (each person's kit number appears in the last column), it is apparent that some have no "definite" unique SNPs (the first two members below, who are brothers) whilst others have up to 5 unique SNPs each (3 people). There are 71 "unique" SNPs in the various tables below, of which 22 are "definite" unique SNPs and the rest are mainly "unlikely" SNPs.

Two things will happen over time: 1) as above, some SNPs will be reclassified (in terms of "definite, probable, possible, and unlikely"), and 2) as more people test on the Big Y (or similar tests), some of these unique SNPs will appear in the new test results and will therefore not be "unique" any more - they will move up from the "Unique SNP" tables below into the "Shared SNP" portion of the tree.

So the Take Home message is this: the unique SNPs will become important over time in identifying further sub-divisions and sub-branches within the Lineage II portion of the human evolutionary tree.

Potentially, because we have 22 "definite" unique SNPs, this will probably translate into 22 new sub-branches. 

And if we add these 22 sub-branches to the 16 sub-branches that would develop from splitting the 4 current SNP blocks within the tree, we should (in time) see an extra 38 sub-branches develop within the Gleason Lineage II SNP-based Mutation History Tree.

However, I suspect that it may be many many more than 38 sub-branches. 

Time will tell.


(click to enlarge)


Maurice Gleeson
Feb 2016




[1] this is not entirely true. Sometimes (rarely?) mutations can occur in the same SNP in entirely different populations, just by chance.

[2] The +, *, **, and *** symbols have slightly different meanings depending on whether the kit is a BigY kit, a FGC kit, or a manual entry of my own.
  • For FTDNA kits, + implies a "PASS" result with just one possible variant, * indicates a "PASS" but with multiple variants, ** indicates "REJECTED" with just a single variant, and *** indicates "REJECTED" with multiple possible variants. The multiple variant mutations tend to fall in repetitive regions.
  • For FGC kits, the meaning of the symbols is the same as it is from the FGC interpretation files. + indicates over 99% likely genuine (95% for INDELs); * over 95% likely genuine (90% for INDELs); ** about 40% likely genuine; *** about 10% likely genuine.
  • Manual entries read directly from a BAM file will be either + indicating positive or * indicating that the data show a mixture of possible variants.