<?xml version="1.0" ?><!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.3 20210610//EN"  "JATS-archivearticle1-mathml3.dtd"><article xmlns:ali="http://www.niso.org/schemas/ali/1.0/" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.3" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="nlm-ta">elife</journal-id>
<journal-id journal-id-type="publisher-id">eLife</journal-id>
<journal-title-group>
<journal-title>eLife</journal-title>
</journal-title-group>
<issn publication-format="electronic" pub-type="epub">2050-084X</issn>
<publisher>
<publisher-name>eLife Sciences Publications, Ltd</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">93258</article-id>
<article-id pub-id-type="doi">10.7554/eLife.93258</article-id>
<article-id pub-id-type="doi" specific-use="version">10.7554/eLife.93258.1</article-id>
<article-version-alternatives>
<article-version article-version-type="publication-state">reviewed preprint</article-version>
<article-version article-version-type="preprint-version">1.2</article-version>
</article-version-alternatives>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Evolutionary Biology</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Analyses of allele age and fitness impact reveal human beneficial alleles to be older than neutral controls</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Pivirotto</surname>
<given-names>Alyssa M.</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Platt</surname>
<given-names>Alexander</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
<xref ref-type="aff" rid="a2">2</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Patel</surname>
<given-names>Ravi</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
<xref ref-type="aff" rid="a3">3</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0002-9918-8212</contrib-id>
<name>
<surname>Kumar</surname>
<given-names>Sudhir</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
<xref ref-type="aff" rid="a3">3</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0001-5358-6488</contrib-id>
<name>
<surname>Hey</surname>
<given-names>Jody</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
<xref ref-type="corresp" rid="cor1">*</xref>
</contrib>
<aff id="a1"><label>1</label><institution>Temple University, Department of Biology</institution>, Philadelphia PA 19122, <country>USA</country></aff>
<aff id="a2"><label>2</label><institution>University of Pennsylvania, Department of Genetics</institution>, Philadelphia PA 19104, <country>USA</country></aff>
<aff id="a3"><label>3</label><institution>Institute for Genomics and Evolutionary Medicine, Temple University</institution>, PA 19122, <country>USA</country></aff>
</contrib-group>
<contrib-group content-type="section">
<contrib contrib-type="editor">
<name>
<surname>Ross-Ibarra</surname>
<given-names>Jeffrey</given-names>
</name>
<role>Reviewing Editor</role>
<aff>
<institution-wrap>
<institution>University of California, Davis</institution>
</institution-wrap>
<city>Davis</city>
<country>United States of America</country>
</aff>
</contrib>
<contrib contrib-type="senior_editor">
<name>
<surname>Perry</surname>
<given-names>George H</given-names>
</name>
<role>Senior Editor</role>
<aff>
<institution-wrap>
<institution>Pennsylvania State University</institution>
</institution-wrap>
<city>University Park</city>
<country>United States of America</country>
</aff>
</contrib>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>*</label>Corresponding Author: Jody Hey <bold>Correspondence:</bold> <email>hey@temple.edu</email></corresp>
</author-notes>
<pub-date date-type="original-publication" iso-8601-date="2024-01-29">
<day>29</day>
<month>01</month>
<year>2024</year>
</pub-date>
<volume>13</volume>
<elocation-id>RP93258</elocation-id>
<history>
<date date-type="sent-for-review" iso-8601-date="2023-10-09">
<day>09</day>
<month>10</month>
<year>2023</year>
</date>
</history>
<pub-history>
<event>
<event-desc>Preprint posted</event-desc>
<date date-type="preprint" iso-8601-date="2023-10-11">
<day>11</day>
<month>10</month>
<year>2023</year>
</date>
<self-uri content-type="preprint" xlink:href="https://doi.org/10.1101/2023.10.09.561569"/>
</event>
</pub-history>
<permissions>
<copyright-statement>© 2024, Pivirotto et al</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Pivirotto et al</copyright-holder>
<ali:free_to_read/>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<ali:license_ref>https://creativecommons.org/licenses/by/4.0/</ali:license_ref>
<license-p>This article is distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License</ext-link>, which permits unrestricted use and redistribution provided that the original author and source are credited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="elife-preprint-93258-v1.pdf"/>
<abstract>
<title>Abstract</title><p>A classic population genetic prediction is that alleles experiencing directional selection should swiftly traverse allele frequency space, leaving detectable reductions in genetic variation in linked regions. However, despite this expectation, identifying clear footprints of beneficial allele passage has proven to be surprisingly challenging. We addressed the basic premise underlying this expectation by estimating the ages of large numbers of beneficial and deleterious alleles in a human population genomic data set. Deleterious alleles were found to be young, on average, given their allele frequency. However, beneficial alleles were older on average than non-coding, non-regulatory alleles of the same frequency. This finding is not consistent with directional selection and instead indicates some type of balancing selection. Among derived beneficial alleles, those fixed in the population show higher local recombination rates than those still segregating, consistent with a model in which new beneficial alleles experience an initial period of balancing selection due to linkage disequilibrium with deleterious recessive alleles. Alleles that ultimately fix following a period of balancing selection will leave a modest ‘soft’ sweep impact on the local variation, consistent with the overall paucity of species-wide ‘hard’ sweeps in human genomes.</p>
</abstract>
<abstract abstract-type="teaser">
<title>Impact Statement</title>
<p>Analyses of allele age and evolutionary impact reveal that beneficial alleles in a human population are often older than neutral controls, suggesting a large role for balancing selection in adaptation.</p>
</abstract>

</article-meta>
<notes>
<notes notes-type="competing-interest-statement">
<title>Competing Interest Statement</title><p>The authors have declared no competing interest.</p></notes>
<fn-group content-type="summary-of-updates">
<title>Summary of Updates:</title>
<fn fn-type="update"><p>Supplementary Information file added</p></fn>
</fn-group>
</notes>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>Evolutionary adaptation depends upon the spread and fixation of beneficial alleles, however some neutral and slightly deleterious alleles also drift to high frequencies and become fixed, and so investigators have long sought ways to distinguish the fixation processes of adaptive alleles from those that are non-adaptive. Most methods are based on the classic population genetic prediction that beneficial alleles should move quickly through the range of allele frequencies (<xref ref-type="bibr" rid="c3">3</xref>, <xref ref-type="bibr" rid="c4">4</xref>) and leave a significant footprint on levels and patterns of linked variation (<xref ref-type="bibr" rid="c5">5</xref>). However, despite evidence that the fixation of beneficial alleles is common (<xref ref-type="bibr" rid="c6">6</xref>–<xref ref-type="bibr" rid="c8">8</xref>), investigators have found few instances where individual fixation events have left a clear footprint (<xref ref-type="bibr" rid="c9">9</xref>–<xref ref-type="bibr" rid="c11">11</xref>). In the human context, this has been particularly puzzling given that other methods suggest that there have been thousands of adaptive amino-acid substitutions in the human lineage since the common ancestor with chimpanzees (<xref ref-type="bibr" rid="c6">6</xref>, <xref ref-type="bibr" rid="c12">12</xref>–<xref ref-type="bibr" rid="c15">15</xref>).</p>
<p>Consequently, much research in recent years has been devoted to understanding the fixation process of beneficial alleles and the kinds of impacts that may be left in contexts of multiple mutations (<xref ref-type="bibr" rid="c16">16</xref>, <xref ref-type="bibr" rid="c17">17</xref>), changing selection coefficients (<xref ref-type="bibr" rid="c18">18</xref>), selection at linked sites (<xref ref-type="bibr" rid="c19">19</xref>), and population structure (<xref ref-type="bibr" rid="c20">20</xref>–<xref ref-type="bibr" rid="c22">22</xref>).</p>
<p>To better understand the allele frequency trajectories of beneficial alleles, we undertook a new kind of analysis that combines two unrelated advances of recent years, one that can identify a large number of segregating beneficial and deleterious alleles, and another that estimates allele age. Our initial goal was to test the fundamental population genetic prediction that alleles under directional selection should be younger, on average, than neutral alleles of the same frequency. This expectation was clearly affirmed for candidate deleterious alleles; however, the analysis revealed a striking pattern in which candidate beneficial alleles are older on average than neutral alleles.</p>
<p>For nonsynonymous single nucleotide polymorphisms (SNPs) in a whole-genome sequencing study of over 3600 individuals from the United Kingdom (<xref ref-type="bibr" rid="c23">23</xref>), we identified candidate alleles under selection using the evolutionary probability (EP) of amino acids residing at each position in 17,209 autosomal genes calculated from a multi-species protein sequence alignment (<xref ref-type="bibr" rid="c24">24</xref>). EP estimates are based on alignments of a large number of vertebrate genomes and do not depend on the alleles currently segregating in a population or their frequency. The use of EP estimates for identifying alleles under selection is well supported by simulation (<xref ref-type="bibr" rid="c25">25</xref>), and they are increasingly used to identify nonsynonymous changes that are candidates for adaptive changes (<xref ref-type="bibr" rid="c26">26</xref>–<xref ref-type="bibr" rid="c30">30</xref>). As shown in <xref rid="fig1" ref-type="fig">Figure 1A</xref>, EP values correlate with allele frequency, with common alleles tending to have higher EP values as expected if high EP alleles are favored by selection more than are low EP alleles.</p>
<fig id="fig1" position="float" orientation="portrait" fig-type="figure">
<label>Figure 1</label>
<caption><p>A. EP for non-synonymous SNPs binned by allele frequency. Both alleles of each SNP are included. Each bin includes a 95% confidence interval on the mean. Sites with higher EP are found at a higher frequency on average while sites with lower EP are found at lower frequencies.</p><p>B. Mean derived-allele frequency binned by ΔEP values. Each bin includes a 95% confidence interval on the mean. Dotted line represents the average frequency of a neutral (non-coding, non-regulatory) site. Higher positive ΔEP bins have a higher frequency on average as expected if these sites are beneficial.</p>
<p>C. EP calculation and age estimation targets for GEVA and <italic>t</italic><sub><italic>c</italic></sub> for a hypothetical site with three copies of the derived allele in a sample of 10 genomes.</p>
</caption>
<graphic xlink:href="561569v2_fig1.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
<p>We rooted non-synonymous variants using the inferred ancestral sequence from Ensembl (<xref ref-type="bibr" rid="c1">1</xref>) and a maximum likelihood estimator. We defined ΔEP as the derived allele EP minus the ancestral allele EP. The large majority of derived alleles are at low frequency, as expected from basic theory (<xref ref-type="bibr" rid="c31">31</xref>), and we observed that mean derived allele frequency increases for sites with higher positive ΔEP (<xref rid="fig1" ref-type="fig">Figure 1B</xref>), as expected if they are favored by natural selection (<xref ref-type="bibr" rid="c32">32</xref>, <xref ref-type="bibr" rid="c33">33</xref>).</p>
<p>To consider the ages of alleles predicted to be under directional selection, we used a large control set of non-coding, non-regulatory SNPs. These will necessarily have experienced similar mutational and recombinational processes, as well as the same demographic history, that non-synonymous SNPs have experienced, and they offer the ideal landscape upon which to inquire of the impact of selection on allele age.</p>
</sec>
<sec id="s2">
<title>Results &amp; Discussion</title>
<sec id="s2a">
<title>Summary of segregating and fixed derived nonsynonymous alleles</title>
<p>With many rooted segregating and fixed SNPs, we can examine some basic expectations of positive and negative directional selection on non-synonymous mutations (<xref rid="tbl1" ref-type="table">Table 1</xref>, Supplemental Table 4). First, if adaptation operates primarily at the margins of optimality, then more non-synonymous variants will be harmful than beneficial, and the magnitude of effect for deleterious mutations should be greater on average than for beneficial mutations (<xref ref-type="bibr" rid="c34">34</xref>). We observe both patterns, with many more negative ΔEP alleles overall, and the mean absolute magnitude of ΔEP is much greater for negative ΔEP SNPs than for positive ΔEP SNPs (0.830 versus 0.274). Comparing fixed and segregating sites, it is expected that derived positive ΔEP alleles with a frequency of 1.0 will have larger ΔEP values than those in which both ancestral and derived alleles occur in the sample, which is confirmed (0.418 for fixed vs. 0.274 for segregating). The same prediction for negative ΔEP SNPs, with fixed alleles having a higher mean value than polymorphic alleles, was also confirmed (-0.685 vs. -0.830).</p>
<table-wrap id="tbl1" orientation="portrait" position="float">
<label>Table 1.</label>
<caption><title>ΔEP measures for fixed and polymorphic alleles.</title></caption>
<graphic xlink:href="561569v2_tbl1.tif" mimetype="image" mime-subtype="tiff"/>
</table-wrap>
</sec>
<sec id="s2b">
<title>Deleterious mutations are younger on average while beneficial mutations are older on average than neutral mutations of the same frequency</title>
<p>Both positively and negatively selected alleles are expected to be younger on average than neutral alleles of the same frequency (<xref ref-type="bibr" rid="c3">3</xref>, <xref ref-type="bibr" rid="c35">35</xref>–<xref ref-type="bibr" rid="c37">37</xref>). We used the Genealogical Estimation of Variant Age (GEVA) method (<xref ref-type="bibr" rid="c38">38</xref>) to estimate the descendent node time, or coalescent time, for genes carrying the derived allele (<xref rid="fig1" ref-type="fig">Figure 1C</xref>). We used RUNTC (<xref ref-type="bibr" rid="c39">39</xref>) to estimate <italic>t</italic><sub><italic>c</italic></sub>, the time of the ancestral node of the edge carrying the mutation (<xref rid="fig1" ref-type="fig">Figure 1C</xref>). Rooted bi-allelic SNPs at non-coding, non-regulatory sites were used for a control set, identified hereafter as “neutral.” The <italic>t</italic><sub><italic>c</italic></sub> estimator is not a function of allele frequency, and GEVA makes only limited use of allele frequency in the setting of priors for the recombinational landscape.</p>
<p>Allele frequency is a strong predictor of allele age, and as expected, the mean derived-allele age rises with frequency for all three classes of SNPs (<xref rid="fig2" ref-type="fig">Figure 2A</xref>). For both positive and negative ΔEP SNPs, an analysis of variance (ANOVA) was conducted to test the hypothesis that selected derived alleles have the same mean age as control SNPs. In both cases the null hypothesis was strongly rejected (p = 4.17x10<sup>-</sup> <sup>12</sup> for negative ΔEP SNPs and 3.04x10<sup>-27</sup> for positive ΔEP SNPs). However, unlike derived negative ΔEP alleles, which were younger on average than control alleles, as predicted, the positive ΔEP SNPs are older on average. Surprisingly, across most frequency intervals, derived positive ΔEP alleles exhibit mean ages thousands of generations older than the neutral control set.</p>
<fig id="fig2" position="float" orientation="portrait" fig-type="figure">
<label>Figure 2</label>
<caption><p>A. Allele age estimates using GEVA by allele frequency, with each frequency bin holding 75,000 neutral sites.</p><p>B. Age rank (GEVA) as a function of ΔEP. Age rank for each derived allele was the rank position of the GEVA estimate in a list of all GEVA ages for neutral alleles with frequency matched derived alleles.</p>
<p>C. Same as B, but for <italic>t</italic><sub><italic>c</italic></sub>.</p></caption>
<graphic xlink:href="561569v2_fig2.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
<p>To isolate the relationship between ΔEP and allele age independently of allele frequency, we placed each allele’s age estimate into an ordered list of ages for neutral alleles of the same frequency. Non- synonymous alleles in the top half of the distribution (ranked higher than 0.5) are thus older than the median age of those neutral alleles. As shown in <xref rid="fig2" ref-type="fig">Figures 2B</xref> and <xref rid="fig2" ref-type="fig">2C</xref>, the ranked ΔEP values show a clear trend, with negative ΔEP values falling consistently below 0.5 (i.e., with ages less than neutral alleles of the same frequency) and positive ΔEP alleles have mean age ranks consistently above 0.5.</p>
<p>Because the set of non-coding, non-regulatory controls necessarily experienced the same demographic context as the selected alleles, explanations of older ages for candidate beneficial alleles that depend upon interactions of selection and demography are largely ruled out, at least for models in which the beneficial alleles are indeed under directional selection. Nor can models in which these alleles are sometimes neutral and sometimes favored help explain the observation, as such alleles would still be expected to be younger on average than our control set. This pattern, in which alleles are maintained longer than alleles that are not subject to selection, is simply not consistent with positive directional selection, but rather suggests some form of balancing selection (<xref ref-type="bibr" rid="c40">40</xref>).</p>
</sec>
<sec id="s2c">
<title>Characterizing old, segregating, positive ΔEP alleles</title>
<p>Overall, a large proportion of positive ΔEP alleles are older than neutral controls. For <italic>t</italic><sub><italic>c</italic></sub> there were 3511 positive ΔEP alleles, 1354 of which had age ranks greater than 0.5 (38.6%). For GEVA there were 1390 positive ΔEP alleles (fewer than for <italic>t</italic><sub><italic>c</italic></sub> as GEVA cannot be applied to alleles that occur only once), 741 of which had age ranks greater than 0.5 (53.3%). We considered the possibility that the elevated ages of segregating positive ΔEP alleles were a kind of sampling artifact, as would occur if they represented the tail of a distribution of ages for all favored alleles, including those that became fixed (which do not appear as SNPs and for which we do not have age estimates). This explanation does not apply to alleles under strong directional selection, for which the mean and variance in sojourn times are low. On the other hand, weakly selected favored alleles will have a large mean and variance in sojourn times (<xref ref-type="bibr" rid="c41">41</xref>), and a large sample of such alleles would have some that, by chance, had been segregating for a long time. However, if the old segregating positive ΔEP alleles were only very weakly favored, and if they constitute the minority of alleles that were held back by the chance effects of genetic drift, then they would make up only a small fraction of all positive ΔEP alleles, including both fixed and segregating. We do not observe this in the data, with segregating alleles constituting a large fraction (0.596, <xref rid="tbl1" ref-type="table">Table 1</xref>) of all positive ΔEP alleles.</p>
<p>Balancing selection can take many forms (<xref ref-type="bibr" rid="c42">42</xref>), but whatever the mode of selection for these alleles, it does not appear to be the kind of long-term balancing selection that causes trans-species polymorphisms like those found in immune-related genes (<xref ref-type="bibr" rid="c43">43</xref>, <xref ref-type="bibr" rid="c44">44</xref>). Of the positive ΔEP alleles, none of the GEVA values, and only 2.5% of the <italic>t</italic><sub><italic>c</italic></sub> values, are over 200,000 generations, which would correspond approximately to the human chimpanzee divergence time, assuming a 29-year generation time (<xref ref-type="bibr" rid="c45">45</xref>).</p>
<p>Most positive ΔEP sites, including those with age ranks greater than 0.5 (i.e., older than neutral alleles of the same frequency) also do not fit a conventional model of balancing selection in that the derived allele frequency is usually low (<xref rid="fig2" ref-type="fig">Figure 2A</xref>, Supplemental Figure 1). For <italic>t</italic><sub><italic>c</italic></sub> the mean frequency of positive ΔEP sites with age ranks greater than 0.5 is 0.039, while for GEVA it is 0.091.</p>
<p>When we seek these alleles in archaic humans, we find that relatively few positive ΔEP alleles identified in the UK10K sample (241; 4.0%) occur in a sample of 4 archaic genomes. The same analysis for negative ΔEP alleles found a smaller proportion of shared alleles (2030; 1.4%), whereas an intermediate value of noncoding sites (401,741; 3.0%) were observed among the sample of archaic genomes. For genomic regions identified as introgressed from archaics, only 13 positive ΔEP alleles (0.2% of all positive ΔEP sites) and 180 negative ΔEP alleles (0.1% of all negative ΔEP sites) were found.</p>
<p>We applied an alternative method for identifying balancing selection to positive ΔEP alleles that is based on the number of nearby polymorphisms that have risen to a similar frequency as the candidate allele (<xref ref-type="bibr" rid="c46">46</xref>). We find that the test statistic, β, is significantly higher for positive ΔEP sites compared to negative ΔEP sites (p-value = 1.588e-6), however the magnitude of these differences is small at just an average β value of 1.09 for positive ΔEP sites and 0.55 for negative ΔEP sites. Because most of the positive ΔEP sites in our study are found at low to moderate frequencies, and because the elevated ages, relative to neutral sites, are on the order of 100’s or 1000’s of generations, it is likely that there has not been sufficient time for genetic drift to bring flanking sites in to the configuration that the β statistic is designed to be sensitive to.</p>
</sec>
<sec id="s2d">
<title>Examination of modes of balancing selection: population structure and overdominance</title>
<p>We observed significant clumping of positive ΔEP SNPs among the genes included in the study. For every autosome, the observed variance in SNP density was significantly greater than that generated by population genetic simulation (Supplemental Table 1). Gene ontology analyses for genes rich in positive ΔEP SNPs revealed enrichment in several categories (Supplemental Table 2), most notably blood coagulation and several disease pathways.</p>
<p>One mechanism that could give rise to new balanced polymorphisms is if the selection regime arose because of the human population structure that favored ancestral alleles in some populations and derived alleles in other populations (as suggested in a recent analysis (<xref ref-type="bibr" rid="c47">47</xref>)). To examine the possibility that population structure is facilitating a large amount of balancing selection, we examined FST in the 1000 genomes data (<xref ref-type="bibr" rid="c48">48</xref>). Analysis of F<sub>ST</sub> values in 1000 Genomes data for alleles from the UK10K samples with positive ΔEP and age ranks greater than 0.5 found no sign that these alleles show greater population structure than control alleles (Supplemental Table 3). In three comparisons, the hypothesis that F<sub>ST</sub> was higher for positive ΔEP alleles that are older than expected could not be rejected by single classification Wilcoxon test in pooled African samples versus pooled European and Asian samples (p = 0.1804), pooled European versus pooled Asian samples (p = 0.5298), and Great Britain sample versus Italian sample (p = 0.7854).</p>
<p>Another possibility is if heterozygous positive ΔEP sites have higher fitness than homozygotes for both the ancestral and the derived alleles. To evaluate this in a way that combined the signal from all positive ΔEP alleles, we asked whether positive ΔEP alleles had higher heterozygote counts than neutral alleles of the same allele frequencies. Analyzing SNPs with at least 100 derived allele copies, we observed equal proportions of positive ΔEP sites with more heterozygotes than the neutral class, compared to fewer; and we found a mean rank for heterozygote count for positive ΔEP sites of 0.501. A one-sided <italic>z-</italic>test of the null hypothesis that the mean rank was equal to or less than 0.5 did not approach statistical significance (p = 0.48). This is consistent with previously published results which failed to find evidence of overdominance at deletion sites thought to be under balancing selection (<xref ref-type="bibr" rid="c49">49</xref>). To assess our ability to detect heterozygote advantage using counts of heterozygotes, a power analysis was conducted using simulations that mirrored the actual data set, assuming genotypes are sampled under heterozygote advantage after selection has acted. The analyses revealed that over a wide range of weak to moderate selection coefficients where the selective advantage is less than 1% (i.e., s &lt; 0.01), that an excess of heterozygotes is unlikely to be detected given the UK10K sample size (Supplementary Table 5).</p>
</sec>
<sec id="s2e">
<title>Models that can account for a period of balancing selection</title>
<p>The absence of very old, derived alleles among positive ΔEP sites suggests that the balancing selection that occurs undergoes a change of character, such that balancing selection occurs for a period of time and is then followed by directional selection or no selection (i.e. genetic drift alone) leading to a loss of one or other of the alleles. If that were not the case, then we would not expect the absence of very old alleles in this data set. To address this, we consider two models that both provide mechanisms for balancing selection and that both predict that balancing selection will be a temporary phase in the process of the fixation of beneficial alleles.</p>
<p>One theory to explain many positive ΔEP alleles with elevated ages includes two selection stages, including first a period of balancing selection under heterozygote advantage, after which positive directional selection carries the allele to fixation. Under this “staggered sweep” model, balancing selection occurs when a favorable allele arises on a chromosome that carries one or more recessive deleterious alleles at nearby locations, and it lasts until recombination moves the allele onto other haplotypes not having linked deleterious alleles (<xref ref-type="bibr" rid="c50">50</xref>). A heterozygote for this chromosomal region is initially favored because of the new allele’s dominance and the harmful allele’s recessivity, such that the net positive selection coefficient on heterozygotes is strong enough to counter the effects of genetic drift. The model is supported by the fact that individual humans, and human populations, carry very large numbers of deleterious alleles, the large majority of which are expected to be mostly recessive in their effects. Considering, for example, just loss-of-function alleles for which diploid European genomes are estimated to carry about 100 (mostly in the heterozygous state), then the odds that a new beneficial mutation arises near to, and in-phase, with a deleterious allele, may be quite high (<xref ref-type="bibr" rid="c51">51</xref>).</p>
<p>Testing the staggered sweep model is difficult because local linkage estimates, as well as <italic>t</italic><sub><italic>c</italic></sub> and GEVA estimates, all depend on a common estimate of the genetic map. However, we can avoid this complication, and partially test the staggered sweep model, by comparing local recombination rates near positive ΔEP alleles that are fixed to those that are segregating. If segregating alleles are under balancing selection because of linkage to deleterious alleles, and the fixed alleles include those that had escaped by recombination, we expect segregating alleles to show lower local recombination rates than fixed positive ΔEP alleles. As predicted, the recombination rates of genomic regions near fixed positive ΔEP alleles were significantly higher than for segregating alleles (Mann Whitney U test p=6.0x10<sup>-19</sup>, <xref rid="fig3" ref-type="fig">Figure 3A</xref>).</p>
<fig id="fig3" position="float" orientation="portrait" fig-type="figure">
<label>Figure 3</label>
<caption><p>A. Mean recombination rate per base per generation as a function of ΔEP for fixed and segregating alleles.</p><p>B. Figurative example of the frequency trajectory of an allele under the staggered sweep (SS) or diploid fisher’s geometric (DFG) model. Both begin with a phase of rising frequency (A) towards a period of equilibrium (B) caused by heterozygote advantage when homozygous genotypes are disfavored, either due to recessive deleterious linked variation (SS) or an overshooting of the optimal phenotype (DFG). Under DFG, variants are ultimately replaced by new mutations that are simply favored. Under SS, alleles eventually cross over onto chromosomes without linked deleterious alleles, and then rise to fixation (C).</p></caption>
<graphic xlink:href="561569v2_fig3.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
<p>Another explanation that also invokes heterozygote advantage is a diploid version of Fisher’s geometric model (denoted hereafter as DFG) in which mutations that carry the phenotype in the direction of the optimum may be favored when heterozygous under codominance and yet disfavored in homozygotes if that phenotype is more extreme and further away from the optimum (<xref ref-type="bibr" rid="c52">52</xref>). Under this model, balancing selection may be a common phase during an adaptive walk toward increasing fitness, with balanced alleles ultimately being lost when new alleles under simple positive directional selection arise and become fixed. The staggered sweep model and the DFG model differ most clearly in that the former has the period of balancing selection as a phase before the fixation of the allele, whereas the latter has the balanced allele being replaced by a new allele that is simply favored by directional selection. The former model predicts that some, perhaps many, selective sweeps are actually ‘soft’ sweeps caused by the fixation of a relatively old allele. In contrast, the DFG model predicts that when a selective sweep occurs, it is a conventional sweep by a new favored allele (i.e., a ‘hard’ sweep). Both models predict partial sweeps around new alleles that arise in a balancing selection fitness scheme (<xref rid="fig3" ref-type="fig">Figure 3B</xref>).</p>
</sec>
<sec id="s2f">
<title>Implications for the adaptation of human populations</title>
<p>We find that the majority of candidate derived beneficial alleles in a human population are segregating, rather than fixed, and yet the mean ages of these alleles are older than those for derived control alleles. These relatively old SNPs do not appear to fit a classical balancing selection model in that most of them are at low frequency and have age estimates almost always less than the age of the hominin branch.</p>
<p>The overall pattern suggests that when fixation of beneficial alleles does occur, it often follows an initial period of balancing selection.</p>
<p>We did not find evidence that ΔEP alleles are maintained due to commonly considered mechanisms of balancing selection such as population structure or heterozygote advantage, although the power to detect these factors was low, unless selection has been quite strong. Instead, we found support for the staggered sweep model in which beneficial alleles arise on the same haplotype as a deleterious mutation which delays them from fixing. Under a staggered sweep model, we predict that there should be differences in recombination rates between segregating and fixed alleles allowing for some alleles to escape selection from nearby deleterious which we find to be true for moderate positive ΔEP sites.</p>
<p>If many beneficial alleles have a lengthy period of balancing selection, before proceeding to fixation, then a significant fraction of adaptive fixations experienced by the human species (not just individual populations) will have occurred as a ‘soft’ sweep rather than a ‘hard’ sweep. This would help explain why there are few unambiguous cases of complete hard sweeps in large population genomic data sets (<xref ref-type="bibr" rid="c9">9</xref>, <xref ref-type="bibr" rid="c11">11</xref>).</p>
<p>An additional implication of these findings is that the process of adaptation by human populations may be slower than basic population genetic models predict. If a significant fraction of ultimately beneficial fixed alleles undergoes a period of balancing selection, then at least at these sites, the process of adaptation is slowed and limited, not for lack of mutation, but rather by the process causing the period of balancing selection.</p>
</sec>
</sec>
<sec id="s3">
<title>Methods</title>
<sec id="s3a">
<title>Evolutionary Probabilities, Allele Frequencies, and Data Filtering</title>
<p>Non-synonymous SNP sites in the UK10K dataset were identified with their corresponding transcript ID using the hg19 RefGene annotations in the UCSC table browser (<xref ref-type="bibr" rid="c53">53</xref>), that are based on NCBI RefSeq annotations (<xref ref-type="bibr" rid="c54">54</xref>), and the UK10K VCF (Variant Calling Format) files (<xref ref-type="bibr" rid="c23">23</xref>). For each two-allele polymorphism, the transcript IDs and site locations were used to retrieve the EP values for both the reference and alternative alleles. EP values were estimated using the method described in previous literature (<xref ref-type="bibr" rid="c24">24</xref>, <xref ref-type="bibr" rid="c55">55</xref>) using posterior probabilities from a multispecies alignment with associated divergence times. Mutations excluded from this dataset include those with un-curated transcript IDs that have not been verified. Frequency data for the reference and alternative allele at each site was extracted directly from the VCF file. Analyses thought to be sensitive to CpG high mutability where limited to SNPs that did not occur as part of a CpG. These included analyses that utilized allele ages (<xref rid="fig2" ref-type="fig">Figures 2A</xref>, 2B, and 2C) as mutation rate was used as a parameter in estimating these values.</p>
</sec>
<sec id="s3b">
<title>Allele Age Estimates</title>
<p>To get approximate allele age estimates, we used both the time of most recent coalescence (<italic>t</italic><sub><italic>c</italic></sub>) estimator (<xref ref-type="bibr" rid="c39">39</xref>) from the Hey Lab and the Genealogical Estimation of Variant Age (GEVA) estimator (<xref ref-type="bibr" rid="c38">38</xref>). To estimate <italic>t</italic><sub><italic>c</italic></sub>, for each of the autosomal chromosome VCF files, first the singletons were phased by placing each singleton on the longer of the two haplotypes. Following this step, the time of coalescence was estimated (runtc.py) using the following parameters: k-range, mutation rate of 1e-8, and recombination map as HapMap Phase II genetic map for hg19 (<xref ref-type="bibr" rid="c56">56</xref>). To obtain GEVA (<xref ref-type="bibr" rid="c38">38</xref>) estimates, the VCF file for each autosomal chromosome was first parsed and converted into a binary file with corresponding marker and site files containing information per variant. GEVA values were obtained for all positive EP SNPs with more than two copies of the derived allele. GEVA estimates were obtained using the default parameters of effective population size of 10000, mutation rate of 1e<sup>-8</sup>, and the provided Hidden Markov Model (HMM) probability files. The output estimated age files were then filtered using the provided program in R (<ext-link ext-link-type="uri" xlink:href="https://github.com/pkalbers/geva">https://github.com/pkalbers/geva</ext-link>).</p>
<p>GEVA estimates were obtained for all positive ΔEP sites in the sampled genes (2729 in total). Because of time constraints large random samples of sites were used for non-coding, non-regulatory sites (71628 in total) and negative ΔEP sites (19053 in total). To generate figures with binned ΔEP values, the number of sampled noncoding, non-regulatory sites range from 800 to 2500 sites with estimated ages. For the negative ΔEP bins have approximately 1000 to 4000 sites with estimated ages, while the positive ΔEP bins have 60 to 600 with estimated ages.</p>
</sec>
<sec id="s3c">
<title>Rooting</title>
<p>Two methods of rooting were used, a parsimony-based approach using Ensembl (<xref ref-type="bibr" rid="c57">57</xref>) and a maximum likelihood approach using RAxML (<xref ref-type="bibr" rid="c58">58</xref>). For the parsimony-based rooting method, estimates of the hg19 ancestral states were retrieved from Ensembl (<xref ref-type="bibr" rid="c1">1</xref>) and included for each position in the dataset. For all analyses of allele age, SNPs were limited to those where the ancestral allele state matched the reference allele. For maximum likelihood rooting a primate alignment was extracted for each RefSeq annotated gene from an Ensembl alignment whole genome alignment (<ext-link ext-link-type="uri" xlink:href="http://ftp.ensembl.org/pub/release-104/maf/ensembl-compara/multiple_alignments/12_primates.epo/12_primates.epo.10_1.maf.gz">http://ftp.ensembl.org/pub/release-104/maf/ensembl-compara/multiple_alignments/12_primates.epo/12_primates.epo.10_1.maf.gz</ext-link>) (<xref ref-type="bibr" rid="c57">57</xref>).</p>
<p>The phylogeny for each gene was estimated using RAxML-NG using the model GTR+Γ. At positions in each gene where there was a non-synonymous mutation in the UK10K dataset, the human sequence base in the alignment was replaced with a missing value, N. Using this newly constructed primate alignment with the modified human sequence to reflect UK10K mutations, RAxML-NG was run again to estimate the base pair values at the base of the edge of the human sequence. The output generated posterior probability estimates for each of the four nucleotides at each non-synonymous SNP site. Using the posterior probabilities, the most likely ancestral state was predicted as the base pair with the highest probability. Downstream analyses were filtered by those sites where a single base pair has a probability above 0.9 indicating a higher certainty for the ancestral state.</p>
</sec>
<sec id="s3d">
<title>Calculating ΔEP</title>
<p>Values of ΔEP were calculated by finding the difference between the derived EP value and the ancestral EP value for a position given an estimated ancestral state for that position. The ΔEP metric indicates the difference from neutrality at a given site between the ancestral and derived allele. Sites where the amino acid mutated from an unlikely state evolutionarily to a more likely state yielded a positive ΔEP value, and in the reverse, sites where the amino acid mutated from a more likely state to less likely state yielded a negative ΔEP value.</p>
</sec>
<sec id="s3e">
<title>Noncoding variants as neutral controls</title>
<p>To account for allele frequency in our analyses of age across the spectrum of ΔEP values, a method to report age in relation to similar frequency control variants was needed. To assess whether an allele was young or old, each allele was compared to a large control set of alleles of the same frequency. For this purpose, we used the ages of noncoding, non-regulatory alleles, treating them as a neutral control set. Candidate SNPs for the control set were first identified from intergenic regions using annotations from SNPeff Human Genome build GRCH37 Ensembl release 75 (<xref ref-type="bibr" rid="c59">59</xref>, <xref ref-type="bibr" rid="c60">60</xref>). This set was then filtered to remove those in regulatory regions, identified as falling into at least one of three data sets available from the UCSC Genome Browser: Candidate cis-Regulatory Elements by ENCODE (<xref ref-type="bibr" rid="c61">61</xref>); RefSeq Functional Elements (<xref ref-type="bibr" rid="c62">62</xref>); and curated regulatory annotations in the ORegAnno database (<xref ref-type="bibr" rid="c63">63</xref>).</p>
<p>Noncoding alleles in non-regulatory regions were assembled into bins of a similar frequency. Of the variants that have identified ancestral states matching the reference allele, noncoding, non-regulatory variants were split into bins of approximately 75,000 variants per frequency bin. At the lower end of frequency bins (k = 1, 2, 3, 4, 5, 6, 7, 8), same k value variants were kept together even if this resulted in bins larger than a size of 75,000 variants. In higher frequency bins, several k values were binned together to yield bins of an approximate size of 75,000 noncoding, nonregulatory variants.</p>
</sec>
<sec id="s3f">
<title>ANOVA</title>
<p>To test the hypotheses that neutral derived allele ages have the same mean as either beneficial or deleterious alleles we used two-way ANOVA, with selected vs control as one effect, and allele frequency bin as a second effect. We first applied the Box-Cox transformation (<xref ref-type="bibr" rid="c64">64</xref>) to GEVA estimates of allele age for each treatment and allele frequency group.</p>
</sec>
<sec id="s3g">
<title>Rank Analysis</title>
<p>To account for differences in allele ages between different frequency bins and to compare variants across the genome, we implemented a ranking system to assign each variant a rank within their own null frequency distribution. Initially, null distributions of noncoding, nonregulatory variant ages were constructed as described above. For each non-synonymous variant remaining in the filtered dataset, the corresponding frequency bin was identified based on the k value of the derived allele at that site. Within the null distribution of ages that correlated to the frequency bin for the focal non-synonymous mutation, the position of the focal mutation’s age within the null distribution was found. Based on that position, the rank within the null distribution was calculated as the position divided by the length of the null distribution (approximately 75,000 variants). This yielded a corresponding rank for each non- synonymous variant based on its own specific null distribution of ages from similar frequency variants.</p>
</sec>
<sec id="s3h">
<title>Recombination Analysis</title>
<p>To identify the changes in recombination across the genome, we found associated recombination rate values for every segregating and fixed non-synonymous site in the UK10K dataset. With all segregating and fixed non-synonymous sites identified using the rooting method described above, the recombination rate at that location was extracted from the genetic map file for the specific demographic in the dataset. In this case, a UK population specific recombination map (<xref ref-type="bibr" rid="c65">65</xref>) was used. With each site’s associated recombination rate, comparisons were made between both fixed and segregating sites across the spectrum of ΔEP values.</p>
</sec>
<sec id="s3i">
<title>F<sub>ST</sub> Analysis</title>
<p>We examined the relationship between F<sub>ST</sub> and ΔEP. In 1000 Genomes data (<xref ref-type="bibr" rid="c48">48</xref>), F<sub>ST</sub> was calculated (<xref ref-type="bibr" rid="c66">66</xref>) for SNPs also found in the UK10K sample for three comparisons: pooled African samples versus pooled European and Asian samples, pooled European versus pooled Asian samples, and Great Britain sample versus Italian sample. Only SNPs with at least 10 copies of the derived allele in the pooled contrast populations were considered. Supplemental Table 3 shows mean F<sub>ST</sub> as a function of ΔEP for each contrast.</p>
<p>To test whether F<sub>ST</sub> was higher for older positive ΔEP SNPs than for control SNPs of the same allele frequencies, the F<sub>ST</sub> for each positive ΔEP SNP with age rank greater than 0.5 was placed in the ranking of F<sub>ST</sub> for all control SNPs of the same derived allele frequency. A single classification Wilcoxon test was conducted on each contrast to test whether there was an excess of positive ΔEP SNPs with F<sub>ST</sub> ranking above 0.5.</p>
</sec>
<sec id="s3j">
<title>Heterozygosity Analysis</title>
<p>A test was conducted for the hypothesis that positive ΔEP SNPs have higher heterozygosity than control SNPs of the same allele frequency. For each positive ΔEP SNP, the rank position of the observed count of the number of heterozygotes was determined by placing the observed count into a sorted list of heterozygote counts for controls SNPs with the same derived allele frequency. In case of ties, the rank position was a random value of all possible ranks with the same heterozygote count. To test the hypothesis that positive ΔEP SNPs have a mean rank above 0.5, a one-sided <italic>z</italic>-test was conducted.</p>
<p>A power analysis was conducted by simulating data sets of the same size and distribution of allele frequencies as the actual data. For a given selection coefficient <italic>s</italic>, where the fitness of a heterozygote is 1+<italic>s</italic>, genotype frequencies were simulated using the observed allele count for each ΔEP SNPs in the data. Heterozygous counts were then placed in corresponding rankings of null distributions of heterozygous counts that were simulated for each of the observed allele frequencies of positive ΔEP SNPs. A <italic>z</italic>-test was conducted for each of 1000 simulated data sets for each selection coefficient. The results are shown in Supplemental Table 5.</p>
</sec>
<sec id="s3k">
<title>Dispersion Analysis</title>
<p>To assess whether positive ΔEP SNPs are evenly distributed among the genes for which we have EP values, we simulated tree-sequence (<xref ref-type="bibr" rid="c67">67</xref>) samples of 7242 UK chromosomes using STDPOPSIM (<xref ref-type="bibr" rid="c68">68</xref>) under an Out-of-Africa model (<xref ref-type="bibr" rid="c69">69</xref>) for each of the autosomes. Then for each autosome mutations were simulated for each gene on that chromosome, using each gene’s actual length and map position, at the same mean density as observed for positive ΔEP SNPS. The variance in simulated density of SNPs was recorded for each of 200 simulations for each autosome.</p>
</sec>
<sec id="s3l">
<title>Gene Ontology Analysis</title>
<p>To test whether positive ΔEP SNPs appeared more often in specific molecular, biological, and cellular classes (GO database released 2022-07-01, DOI: 10.5281/zenodo.6799722), PANTHER pathways (<xref ref-type="bibr" rid="c70">70</xref>) and protein classes (version 17.0, released 2022-02-22), and Reactome Pathways (Reactome database version 77, released 2021-10-01), a PANTHER Overrepresentation Test (Release 20221013) was used (<xref ref-type="bibr" rid="c71">71</xref>, <xref ref-type="bibr" rid="c72">72</xref>). The analyzed set of genes were identified by counting the number of positive ΔEP SNPs per gene. The number of positive ΔEP SNPs was normalized by gene length, and all genes with more than one positive ΔEP were retained. A final subset of 73 genes were used in the PANTHER GO term analysis.</p>
<p>For the reference list, the gene database for Homo sapiens was used. Analyses were conducted with a Fisher’s Exact test with a False Discovery Rate correction. Results are detailed in Supplemental Table 2.</p>
</sec>
<sec id="s3m">
<title>Comparison to Archaic Genomes</title>
<p>In order to identify whether a large proportion of our sites of interest arose prior to the speciation between modern humans and archaic humans, we examined for each site whether it was also present in any one of four archaic genomes (<xref ref-type="bibr" rid="c73">73</xref>–<xref ref-type="bibr" rid="c76">76</xref>). For each category: nonsynonymous – ΔEP, nonsynonymous + ΔEP, and neutral noncoding sites, the number of shared loci with at least one archaic genome is reported along with percent of shared sites over the number of all sites in that category.</p>
<p>Not only was there interest in knowing whether these sites arose prior to the speciation event, but some subset of these sites potentially could be found in both modern human genomes and archaic human genomes due to gene flow between the two species. Sites were identified as appearing in introgression regions based on S* values generated from the CEU dataset from 1000 Genomes (<xref ref-type="bibr" rid="c77">77</xref>) (available at <ext-link ext-link-type="uri" xlink:href="https://data.mendeley.com/datasets/y7hyt83vxr/1">https://data.mendeley.com/datasets/y7hyt83vxr/1</ext-link>). Sites annotated as matching in either Neanderthal or Denisovan would be included as introgression sites for our analysis.</p>
</sec>
<sec id="s3n">
<title>β (<sup><xref ref-type="bibr" rid="c2">2</xref></sup>) Values</title>
<p>For β (<sup><xref ref-type="bibr" rid="c2">2</xref></sup>) scores (<xref ref-type="bibr" rid="c46">46</xref>), the CEU standardized scores generated from 1000 Genomes data was used (available at <ext-link ext-link-type="uri" xlink:href="https://zenodo.org/record/7842447">https://zenodo.org/record/7842447</ext-link>). For each site in our analysis, we identified from this published dataset the Beta2 score if available. A Mann-Whitney U test was done to analyze the difference between the Beta2 values of the – ΔEP and + ΔEP distributions.</p>
</sec>
</sec>
<sec id="d1e934" sec-type="supplementary-material">
<title>Supporting information</title>
<supplementary-material id="d1e1022">
<label>Supplementary Information</label>
<media xlink:href="supplements/561569_file02.pdf"/>
</supplementary-material>
</sec>
</body>
<back>
<ack>
<title>Acknowledgments</title>
<p>This research was supported in part by NIH grants R01GM144468-01 to J. Hey and R35GM139540-02 to S. Kumar. A. Platt was partially funded by N.I.H. grant R35 GM134957-01 and American Diabetes Association Pathway to Stop Diabetes grant #1-19-VSN-02. This research includes calculations carried out on HPC (High Performance Computing) resources supported in part by the National Science Foundation through major research instrumentation grant number 1625061 and by the US Army Research Laboratory under contract number W911NF-16-2-0189.</p>
</ack>
<sec id="s4">
<title>Data Availability</title>
<p>Tables of detailed information for nonsynonymous and noncoding variants, as well as a list of primary mRNA isoforms (in the form of RefSeq IDs) used to retrieve EP values, are available at <ext-link ext-link-type="uri" xlink:href="https://bio.cst.temple.edu/~tuf29449/nolinks/Pivirotto_Balancing_Selection_info.zip">https://bio.cst.temple.edu/~tuf29449/nolinks/Pivirotto_Balancing_Selection_info.zip</ext-link>.</p>
</sec>
<sec id="s5">
<title>Author Contributions</title>
<p>AMP, SK, AP, and JH developed the idea for the study. RP contributed evolutionary probability values. AMP and JH conducted the study, including writing scripts and conducted the analyses. AMP and JH drafted the paper, with comments and suggestions from AP and SK.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="c1"><label>1.</label><mixed-citation publication-type="journal"><string-name><surname>Herrero</surname> <given-names>J</given-names></string-name>, <string-name><surname>Muffato</surname> <given-names>M</given-names></string-name>, <string-name><surname>Beal</surname> <given-names>K</given-names></string-name>, <string-name><surname>Fitzgerald</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gordon</surname> <given-names>L</given-names></string-name>, <string-name><surname>Pignatelli</surname> <given-names>M</given-names></string-name>, <etal>et al.</etal> <article-title>Ensembl comparative genomics resources</article-title>. <source>Database</source>. <volume>2016</volume>;<year>2016</year>.</mixed-citation></ref>
<ref id="c2"><label>2.</label><mixed-citation publication-type="book"><string-name><surname>Efron</surname> <given-names>B</given-names></string-name>, <source>Tibshirani RJ</source>. <publisher-loc>An introduction to the bootstrap</publisher-loc>: <publisher-name>CRC press</publisher-name>; <year>1994</year>.</mixed-citation></ref>
<ref id="c3"><label>3.</label><mixed-citation publication-type="journal"><string-name><surname>Maruyama</surname> <given-names>T</given-names></string-name>. <article-title>The age of an allele in a finite population</article-title>. <source>Genet Res</source>. <year>1974</year>;<volume>23</volume>(<issue>02</issue>):<fpage>137</fpage>–<lpage>43</lpage>.</mixed-citation></ref>
<ref id="c4"><label>4.</label><mixed-citation publication-type="journal"><string-name><surname>Kimura</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ohta</surname> <given-names>T</given-names></string-name>. <article-title>The average number of generations until fixation of a mutant gene in a finite population</article-title>. <source>Genetics</source>. <year>1969</year>;<volume>61</volume>:<fpage>763</fpage>–<lpage>71</lpage>.</mixed-citation></ref>
<ref id="c5"><label>5.</label><mixed-citation publication-type="journal"><string-name><given-names>Maynard</given-names> <surname>Smith J</surname></string-name>, <string-name><surname>Haigh</surname> <given-names>J</given-names></string-name>. <article-title>The hitch-hiking effect of a favourable gene</article-title>. <source>Genet Res</source>. <year>1974</year>;<volume>23</volume>:<fpage>23</fpage>–<lpage>35</lpage>.</mixed-citation></ref>
<ref id="c6"><label>6.</label><mixed-citation publication-type="journal"><string-name><surname>Uricchio</surname> <given-names>LH</given-names></string-name>, <string-name><surname>Petrov</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Enard</surname> <given-names>D</given-names></string-name>. <article-title>Exploiting selection at linked sites to infer the rate and strength of adaptation</article-title>. <source>Nature Ecology &amp; Evolution</source>. <year>2019</year>.</mixed-citation></ref>
<ref id="c7"><label>7.</label><mixed-citation publication-type="journal"><string-name><surname>Galtier</surname> <given-names>N</given-names></string-name>. <article-title>Adaptive Protein Evolution in Animals and the Effective Population Size Hypothesis</article-title>. <source>PLoS Genetics</source>. <year>2016</year>;<volume>12</volume>(<issue>1</issue>):<fpage>e1005774</fpage>.</mixed-citation></ref>
<ref id="c8"><label>8.</label><mixed-citation publication-type="journal"><string-name><surname>Enard</surname> <given-names>D</given-names></string-name>, <string-name><surname>Messer</surname> <given-names>PW</given-names></string-name>, <string-name><surname>Petrov</surname> <given-names>DA</given-names></string-name>. <article-title>Genome-wide signals of positive selection in human evolution</article-title>. <source>Genome Research</source>. <year>2014</year>;<volume>24</volume>(<issue>6</issue>):<fpage>885</fpage>–<lpage>95</lpage>.</mixed-citation></ref>
<ref id="c9"><label>9.</label><mixed-citation publication-type="journal"><string-name><surname>Hernandez</surname> <given-names>RD</given-names></string-name>, <string-name><surname>Kelley</surname> <given-names>JL</given-names></string-name>, <string-name><surname>Elyashiv</surname> <given-names>E</given-names></string-name>, <string-name><surname>Melton</surname> <given-names>SC</given-names></string-name>, <string-name><surname>Auton</surname> <given-names>A</given-names></string-name>, <string-name><surname>McVean</surname> <given-names>G</given-names></string-name>, <etal>et al.</etal> <article-title>Classic Selective Sweeps Were Rare in Recent Human Evolution</article-title>. <source>Science</source>. <year>2011</year>;<volume>331</volume>(6019):<fpage>920</fpage>-4.</mixed-citation></ref>
<ref id="c10"><label>10.</label><mixed-citation publication-type="journal"><string-name><surname>Schrider</surname> <given-names>DR</given-names></string-name>, <string-name><surname>Kern</surname> <given-names>AD</given-names></string-name>. <article-title>Soft sweeps are the dominant mode of adaptation in the human genome</article-title>. <source>Mol Biol Evol</source>. <year>2017</year>;<volume>34</volume>(<issue>8</issue>):<fpage>1863</fpage>–<lpage>77</lpage>.</mixed-citation></ref>
<ref id="c11"><label>11.</label><mixed-citation publication-type="journal"><string-name><surname>Coop</surname> <given-names>G</given-names></string-name>, <string-name><surname>Pickrell</surname> <given-names>JK</given-names></string-name>, <string-name><surname>Novembre</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kudaravalli</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Absher</surname> <given-names>D</given-names></string-name>, <etal>et al.</etal> <article-title>The Role of Geography in Human Adaptation</article-title>. <source>PLoS Genet</source>. <year>2009</year>;<volume>5</volume>(<issue>6</issue>):<fpage>e1000500</fpage>.</mixed-citation></ref>
<ref id="c12"><label>12.</label><mixed-citation publication-type="journal"><collab>12. Consortium CSaA.</collab> <source>Initial sequence of the chimpanzee genome and comparison with the human genome</source>. <year>2005</year>;<volume>437</volume>(<issue>7055</issue>):<fpage>69</fpage>–<lpage>87</lpage>.</mixed-citation></ref>
<ref id="c13"><label>13.</label><mixed-citation publication-type="journal"><string-name><surname>Zhen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huber</surname> <given-names>CD</given-names></string-name>, <string-name><surname>Davies</surname> <given-names>RW</given-names></string-name>, <string-name><surname>Lohmueller</surname> <given-names>KE</given-names></string-name>. <article-title>Greater strength of selection and higher proportion of beneficial amino acid changing mutations in humans compared with mice and Drosophila melanogaster</article-title>. <source>Genome Res</source>. <year>2021</year>;<volume>31</volume>(<issue>1</issue>):<fpage>110</fpage>–<lpage>20</lpage>.</mixed-citation></ref>
<ref id="c14"><label>14.</label><mixed-citation publication-type="journal"><string-name><surname>Huber</surname> <given-names>CD</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>BY</given-names></string-name>, <string-name><surname>Marsden</surname> <given-names>CD</given-names></string-name>, <string-name><surname>Lohmueller</surname> <given-names>KE</given-names></string-name>. <article-title>Determining the factors driving selective effects of new nonsynonymous mutations</article-title>. <source>Proceedings of the National Academy of Sciences</source>. <year>2017</year>;<volume>114</volume>(<issue>17</issue>):<fpage>4465</fpage>–<lpage>70</lpage>.</mixed-citation></ref>
<ref id="c15"><label>15.</label><mixed-citation publication-type="journal"><string-name><surname>Boyko</surname> <given-names>AR</given-names></string-name>, <string-name><surname>Williamson</surname> <given-names>SH</given-names></string-name>, <string-name><surname>Indap</surname> <given-names>AR</given-names></string-name>, <string-name><surname>Degenhardt</surname> <given-names>JD</given-names></string-name>, <string-name><surname>Hernandez</surname> <given-names>RD</given-names></string-name>, <string-name><surname>Lohmueller</surname> <given-names>KE</given-names></string-name>, <etal>et al.</etal> <article-title>Assessing the evolutionary impact of amino acid mutations in the human genome</article-title>. <source>PLoS Genet</source>. <year>2008</year>;<volume>4</volume>(<issue>5</issue>):<fpage>e1000083</fpage>.</mixed-citation></ref>
<ref id="c16"><label>16.</label><mixed-citation publication-type="journal"><string-name><surname>Garud</surname> <given-names>NR</given-names></string-name>, <string-name><surname>Messer</surname> <given-names>PW</given-names></string-name>, <string-name><surname>Petrov</surname> <given-names>DA</given-names></string-name>. <article-title>Detection of hard and soft selective sweeps from Drosophila melanogaster population genomic data</article-title>. <source>PLoS Genetics</source>. <year>2021</year>;<volume>17</volume>(<issue>2</issue>):<fpage>e1009373</fpage>.</mixed-citation></ref>
<ref id="c17"><label>17.</label><mixed-citation publication-type="journal"><string-name><surname>Harris</surname> <given-names>RB</given-names></string-name>, <string-name><surname>Sackman</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jensen</surname> <given-names>JD</given-names></string-name>. <article-title>On the unfounded enthusiasm for soft selective sweeps II: Examining recent evidence from humans, flies, and viruses</article-title>. <source>PLoS Genetics</source>. <year>2018</year>;<volume>14</volume>(<issue>12</issue>):<fpage>e1007859</fpage>.</mixed-citation></ref>
<ref id="c18"><label>18.</label><mixed-citation publication-type="journal"><string-name><surname>McCoy</surname> <given-names>RC</given-names></string-name>, <string-name><surname>Akey</surname> <given-names>JM</given-names></string-name>. <article-title>Selection plays the hand it was dealt: evidence that human adaptation commonly targets standing genetic variation</article-title>. <source>Genome biology</source>. <year>2017</year>;<volume>18</volume>(<issue>1</issue>):<fpage>1</fpage>–<lpage>4</lpage>.</mixed-citation></ref>
<ref id="c19"><label>19.</label><mixed-citation publication-type="journal"><string-name><surname>Charlesworth</surname> <given-names>B</given-names></string-name>, <string-name><surname>Jensen</surname> <given-names>JD</given-names></string-name>. <article-title>Effects of Selection at Linked Sites on Patterns of Genetic Variability</article-title>. <source>Annual Review of Ecology, Evolution, and Systematics</source>. <year>2021</year>;<volume>52</volume>(<issue>1</issue>):<fpage>177</fpage>–<lpage>97</lpage>.</mixed-citation></ref>
<ref id="c20"><label>20.</label><mixed-citation publication-type="journal"><string-name><surname>Souilmi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tobler</surname> <given-names>R</given-names></string-name>, <string-name><surname>Johar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Williams</surname> <given-names>M</given-names></string-name>, <string-name><surname>Grey</surname> <given-names>ST</given-names></string-name>, <string-name><surname>Schmidt</surname> <given-names>J</given-names></string-name>, <etal>et al.</etal> <article-title>Admixture has obscured signals of historical hard sweeps in humans</article-title>. <source>Nature Ecology &amp; Evolution</source>. <year>2022</year>:<fpage>1</fpage>–<lpage>13</lpage>.</mixed-citation></ref>
<ref id="c21"><label>21.</label><mixed-citation publication-type="journal"><string-name><surname>Novembre</surname> <given-names>J</given-names></string-name>, <string-name><surname>Galvani</surname> <given-names>AP</given-names></string-name>, <string-name><surname>Slatkin</surname> <given-names>M</given-names></string-name>. <article-title>The geographic spread of the CCR5 Δ32 HIV-resistance allele</article-title>. <source>PLoS Biology</source>. <year>2005</year>;<volume>3</volume>(<issue>11</issue>):<fpage>e339</fpage>.</mixed-citation></ref>
<ref id="c22"><label>22.</label><mixed-citation publication-type="journal"><string-name><surname>Muktupavela</surname> <given-names>RA</given-names></string-name>, <string-name><surname>Petr</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ségurel</surname> <given-names>L</given-names></string-name>, <string-name><surname>Korneliussen</surname> <given-names>T</given-names></string-name>, <string-name><surname>Novembre</surname> <given-names>J</given-names></string-name>, <string-name><surname>Racimo</surname> <given-names>F</given-names></string-name>. <article-title>Modeling the spatiotemporal spread of beneficial alleles using ancient genomes</article-title>. <source>Elife</source>. <year>2022</year>;<volume>11</volume>:<fpage>e73767</fpage>.</mixed-citation></ref>
<ref id="c23"><label>23.</label><mixed-citation publication-type="journal"><string-name><surname>Consortium</surname> <given-names>TUK</given-names></string-name>. <article-title>The UK10K project identifies rare variants in health and disease</article-title>. <source>Nature</source>. <year>2015</year>;<volume>526</volume>(7571):<fpage>82</fpage>–<lpage>90</lpage>.</mixed-citation></ref>
<ref id="c24"><label>24.</label><mixed-citation publication-type="journal"><string-name><surname>Patel</surname> <given-names>R</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>S</given-names></string-name>. <article-title>On estimating evolutionary probabilities of population variants</article-title>. <source>BMC Evolutionary Biology</source>. <year>2019</year>;<volume>19</volume>(<issue>1</issue>):<fpage>1</fpage>–<lpage>14</lpage>.</mixed-citation></ref>
<ref id="c25"><label>25.</label><mixed-citation publication-type="journal"><string-name><surname>Patel</surname> <given-names>R</given-names></string-name>, <string-name><surname>Scheinfeldt</surname> <given-names>LB</given-names></string-name>, <string-name><surname>Sanderford</surname> <given-names>MD</given-names></string-name>, <string-name><surname>Lanham</surname> <given-names>TR</given-names></string-name>, <string-name><surname>Tamura</surname> <given-names>K</given-names></string-name>, <string-name><surname>Platt</surname> <given-names>A</given-names></string-name>, <etal>et al.</etal> <article-title>Adaptive landscape of protein variation in human exomes</article-title>. <source>Mol Biol Evol</source>. <year>2018</year>;<volume>35</volume>(<issue>8</issue>):<fpage>2015</fpage>–<lpage>25</lpage>.</mixed-citation></ref>
<ref id="c26"><label>26.</label><mixed-citation publication-type="journal"><string-name><surname>Pyott</surname> <given-names>SJ</given-names></string-name>, <string-name><surname>van Tuinen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Screven</surname> <given-names>LA</given-names></string-name>, <string-name><surname>Schrode</surname> <given-names>KM</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>J-P</given-names></string-name>, <string-name><surname>Barone</surname> <given-names>CM</given-names></string-name>, <etal>et al.</etal> <article-title>Functional, morphological, and evolutionary characterization of hearing in subterranean, eusocial African mole-rats</article-title>. <source>Curr Biol</source>. <year>2020</year>;<volume>30</volume>(<issue>22</issue>):<fpage>4329</fpage>–<lpage>41</lpage>. e4.</mixed-citation></ref>
<ref id="c27"><label>27.</label><mixed-citation publication-type="journal"><string-name><surname>Dolatyabi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Peighambari</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Razmyar</surname> <given-names>J</given-names></string-name>. <article-title>Molecular detection and analysis of beak and feather disease viruses in Iran</article-title>. <source>Frontiers in Veterinary Science</source>. <year>2022</year>;<volume>9</volume>.</mixed-citation></ref>
<ref id="c28"><label>28.</label><mixed-citation publication-type="journal"><string-name><surname>Xu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Kosoy</surname> <given-names>R</given-names></string-name>, <string-name><surname>Shameer</surname> <given-names>K</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Readhead</surname> <given-names>B</given-names></string-name>, <etal>et al.</etal> <article-title>Genome-wide analysis indicates association between heterozygote advantage and healthy aging in humans</article-title>. <source>BMC genetics</source>. <year>2019</year>;<volume>20</volume>(<issue>1</issue>):<fpage>1</fpage>–<lpage>14</lpage>.</mixed-citation></ref>
<ref id="c29"><label>29.</label><mixed-citation publication-type="journal"><string-name><surname>Tian</surname> <given-names>R</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Etheridge</surname> <given-names>TH</given-names></string-name>, <string-name><surname>Deshmukh</surname> <given-names>H</given-names></string-name>, <string-name><surname>Gulick</surname> <given-names>D</given-names></string-name>, <string-name><surname>Gibson</surname> <given-names>G</given-names></string-name>, <etal>et al.</etal> <article-title>Pitfalls in single clone CRISPR-Cas9 mutagenesis to fine-map regulatory intervals</article-title>. <source>Genes</source>. <year>2020</year>;<volume>11</volume>(<issue>5</issue>):<fpage>504</fpage>.</mixed-citation></ref>
<ref id="c30"><label>30.</label><mixed-citation publication-type="journal"><string-name><surname>Ose</surname> <given-names>NJ</given-names></string-name>, <string-name><surname>Campitelli</surname> <given-names>P</given-names></string-name>, <string-name><surname>Patel</surname> <given-names>R</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ozkan</surname> <given-names>SB</given-names></string-name>. <article-title>Protein dynamics provide mechanistic insights about epistasis among common missense polymorphisms</article-title>. <source>Biophysical journal</source>. <year>2023</year>.</mixed-citation></ref>
<ref id="c31"><label>31.</label><mixed-citation publication-type="journal"><string-name><surname>Wright</surname> <given-names>S</given-names></string-name>. <article-title>The distribution of gene frequencies in populations</article-title>. <source>Proc Natl Acad Sci U S A</source>. <year>1937</year>;<volume>23</volume>:<fpage>307</fpage>–<lpage>20</lpage>.</mixed-citation></ref>
<ref id="c32"><label>32.</label><mixed-citation publication-type="journal"><string-name><surname>Wright</surname> <given-names>S</given-names></string-name>. <article-title>The Distribution of Gene Frequencies Under Irreversible Mutation</article-title>. <source>Proc Natl Acad Sci</source>. <year>1938</year>;<volume>24</volume>(<issue>7</issue>):<fpage>253</fpage>–<lpage>9</lpage>.</mixed-citation></ref>
<ref id="c33"><label>33.</label><mixed-citation publication-type="journal"><string-name><surname>Kimura</surname> <given-names>M</given-names></string-name>. <article-title>Genetic variability maintained in a finite population due to mutational production of neutral and nearly neutral isoalleles</article-title>. <source>Genet Res</source>. <year>1968</year>;<volume>11</volume>:<fpage>247</fpage>–<lpage>69</lpage>.</mixed-citation></ref>
<ref id="c34"><label>34.</label><mixed-citation publication-type="book"><string-name><surname>Fisher</surname> <given-names>RA</given-names></string-name>. <source>The genetical theory of natural selection</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Clarenson Press</publisher-name>; <year>1930</year>.</mixed-citation></ref>
<ref id="c35"><label>35.</label><mixed-citation publication-type="journal"><string-name><surname>Kimura</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ohta</surname> <given-names>T</given-names></string-name>. <article-title>The age of a neutral mutant persisting in a finite population</article-title>. <source>Genetics</source>. <year>1973</year>;<volume>75</volume>:<fpage>199</fpage>–<lpage>212</lpage>.</mixed-citation></ref>
<ref id="c36"><label>36.</label><mixed-citation publication-type="journal"><string-name><surname>Slatkin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Rannala</surname> <given-names>B</given-names></string-name>. <article-title>Estimating Allele Age</article-title>. <source>Annual Review of Genomics and Human Genetics</source>. <year>2000</year>;<volume>1</volume>(<issue>1</issue>):<fpage>225</fpage>–<lpage>49</lpage>.</mixed-citation></ref>
<ref id="c37"><label>37.</label><mixed-citation publication-type="journal"><string-name><surname>Kiezun</surname> <given-names>A</given-names></string-name>, <string-name><surname>Pulit</surname> <given-names>SL</given-names></string-name>, <string-name><surname>Francioli</surname> <given-names>LC</given-names></string-name>, <string-name><surname>van Dijk</surname> <given-names>F</given-names></string-name>, <string-name><surname>Swertz</surname> <given-names>M</given-names></string-name>, <string-name><surname>Boomsma</surname> <given-names>DI</given-names></string-name>, <etal>et al.</etal> <article-title>Deleterious Alleles in the Human Genome Are on Average Younger Than Neutral Alleles of the Same Frequency</article-title>. <source>PLoS Genet</source>. <year>2013</year>;<volume>9</volume>(<issue>2</issue>):<fpage>e1003301</fpage>.</mixed-citation></ref>
<ref id="c38"><label>38.</label><mixed-citation publication-type="journal"><string-name><surname>Albers</surname> <given-names>PK</given-names></string-name>, <string-name><surname>McVean</surname> <given-names>G</given-names></string-name>. <article-title>Dating genomic variants and shared ancestry in population-scale sequencing data</article-title>. <source>PLoS Biology</source>. <year>2020</year>;<volume>18</volume>(<issue>1</issue>):<fpage>e3000586</fpage>.</mixed-citation></ref>
<ref id="c39"><label>39.</label><mixed-citation publication-type="journal"><string-name><surname>Platt</surname> <given-names>A</given-names></string-name>, <string-name><surname>Pivirotto</surname> <given-names>A</given-names></string-name>, <string-name><surname>Knoblauch</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hey</surname> <given-names>J</given-names></string-name>. <article-title>An estimator of first coalescent time reveals selection on young variants and large heterogeneity in rare allele ages among human populations</article-title>. <source>PLoS Genetics</source>. <year>2019</year>;<volume>15</volume>(<issue>8</issue>):<fpage>e1008340</fpage>.</mixed-citation></ref>
<ref id="c40"><label>40.</label><mixed-citation publication-type="journal"><string-name><given-names>Dobzhansky T.</given-names> <surname>Mendelism</surname></string-name>, <article-title>Darwinism, and evolutionism</article-title>. <source>Proc Am Philos Soc</source>. <year>1965</year>;<volume>109</volume>(<issue>4</issue>):<fpage>205</fpage>–<lpage>15</lpage>.</mixed-citation></ref>
<ref id="c41"><label>41.</label><mixed-citation publication-type="journal"><string-name><surname>De Sanctis</surname> <given-names>B</given-names></string-name>, <string-name><surname>Krukov</surname> <given-names>I</given-names></string-name>, <string-name><surname>de Koning</surname> <given-names>A</given-names></string-name>. <article-title>Allele age under non-classical assumptions is clarified by an exact computational Markov chain approach</article-title>. <source>Scientific reports</source>. <year>2017</year>;<volume>7</volume>(<issue>1</issue>):<fpage>1</fpage>–<lpage>11</lpage>.</mixed-citation></ref>
<ref id="c42"><label>42.</label><mixed-citation publication-type="book"><string-name><surname>Dobzhansky</surname> <given-names>T.</given-names></string-name> <chapter-title>Genetics of the Evolutionary Process</chapter-title>: <publisher-name>Columbia University Press</publisher-name>; <year>1971</year>.</mixed-citation></ref>
<ref id="c43"><label>43.</label><mixed-citation publication-type="journal"><string-name><surname>Leffler</surname> <given-names>EM</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Pfeifer</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ségurel</surname> <given-names>L</given-names></string-name>, <string-name><surname>Auton</surname> <given-names>A</given-names></string-name>, <string-name><surname>Venn</surname> <given-names>O</given-names></string-name>, <etal>et al.</etal> <article-title>Multiple instances of ancient balancing selection shared between humans and chimpanzees</article-title>. <source>Science</source>. <year>2013</year>;<volume>339</volume>(6127):<fpage>1578</fpage>-82.</mixed-citation></ref>
<ref id="c44"><label>44.</label><mixed-citation publication-type="journal"><string-name><surname>Bitarello</surname> <given-names>BD</given-names></string-name>, <string-name><surname>de Filippo</surname> <given-names>C</given-names></string-name>, <string-name><surname>Teixeira</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Schmidt</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Kleinert</surname> <given-names>P</given-names></string-name>, <string-name><surname>Meyer</surname> <given-names>D</given-names></string-name>, <etal>et al.</etal> <article-title>Signatures of long- term balancing selection in human genomes</article-title>. <source>Genome biology and evolution</source>. <year>2018</year>;<volume>10</volume>(<issue>3</issue>):<fpage>939</fpage>–<lpage>55</lpage>.</mixed-citation></ref>
<ref id="c45"><label>45.</label><mixed-citation publication-type="journal"><string-name><surname>Fenner</surname> <given-names>JN</given-names></string-name>. <article-title>Cross-cultural estimation of the human generation interval for use in genetics-based population divergence studies</article-title>. <source>Am J Phys Anthrop</source>. <year>2005</year>;<volume>128</volume>(<issue>2</issue>):<fpage>415</fpage>–<lpage>23</lpage>.</mixed-citation></ref>
<ref id="c46"><label>46.</label><mixed-citation publication-type="journal"><string-name><surname>Siewert</surname> <given-names>KM</given-names></string-name>, <string-name><surname>Voight</surname> <given-names>BF</given-names></string-name>. <article-title>BetaScan2: Standardized statistics to detect balancing selection utilizing substitution data</article-title>. <source>Genome Biology and Evolution</source>. <year>2020</year>;<volume>12</volume>(<issue>2</issue>):<fpage>3873</fpage>–<lpage>7</lpage>.</mixed-citation></ref>
<ref id="c47"><label>47.</label><mixed-citation publication-type="journal"><string-name><surname>Soni</surname> <given-names>V</given-names></string-name>, <string-name><surname>Vos</surname> <given-names>M</given-names></string-name>, <string-name><surname>Eyre-Walker</surname> <given-names>A</given-names></string-name>. <article-title>A new test suggests hundreds of amino acid polymorphisms in humans are subject to balancing selection</article-title>. <source>PLoS Biology</source>. <year>2022</year>;<volume>20</volume>(<issue>6</issue>):<fpage>e3001645</fpage>.</mixed-citation></ref>
<ref id="c48"><label>48.</label><mixed-citation publication-type="journal"><collab>1000 Genomes Project Consortium</collab>. <article-title>A global reference for human genetic variation</article-title>. <source>Nature</source>. <year>2015</year>;<volume>526</volume>(<issue>7571</issue>):<fpage>68</fpage>-<lpage>74</lpage>.</mixed-citation></ref>
<ref id="c49"><label>49.</label><mixed-citation publication-type="journal"><string-name><surname>Aqil</surname> <given-names>A</given-names></string-name>, <string-name><surname>Speidel</surname> <given-names>L</given-names></string-name>, <string-name><surname>Pavlidis</surname> <given-names>P</given-names></string-name>, <string-name><surname>Gokcumen</surname> <given-names>O</given-names></string-name>. <article-title>Balancing selection on genomic deletion polymorphisms in humans</article-title>. <source>Elife</source>. <year>2023</year>;<volume>12</volume>:<fpage>e79111</fpage>.</mixed-citation></ref>
<ref id="c50"><label>50.</label><mixed-citation publication-type="journal"><string-name><surname>Assaf</surname> <given-names>ZJ</given-names></string-name>, <string-name><surname>Petrov</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Blundell</surname> <given-names>JR</given-names></string-name>. <article-title>Obstruction of adaptation in diploids by recessive, strongly deleterious alleles</article-title>. <source>Proceedings of the National Academy of Sciences</source>. <year>2015</year>;<volume>112</volume>(<issue>20</issue>):<fpage>E2658</fpage>–<lpage>E66</lpage>.</mixed-citation></ref>
<ref id="c51"><label>51.</label><mixed-citation publication-type="journal"><string-name><surname>Henn</surname> <given-names>BM</given-names></string-name>, <string-name><surname>Botigué</surname> <given-names>LR</given-names></string-name>, <string-name><surname>Bustamante</surname> <given-names>CD</given-names></string-name>, <string-name><surname>Clark</surname> <given-names>AG</given-names></string-name>, <string-name><surname>Gravel</surname> <given-names>S</given-names></string-name>. <article-title>Estimating Mutation Load in Human Genomes</article-title>. <source>Nature reviews Genetics</source>. <year>2015</year>;<volume>16</volume>(<issue>6</issue>):<fpage>333</fpage>–<lpage>43</lpage>.</mixed-citation></ref>
<ref id="c52"><label>52.</label><mixed-citation publication-type="journal"><string-name><surname>Sellis</surname> <given-names>D</given-names></string-name>, <string-name><surname>Callahan</surname> <given-names>BJ</given-names></string-name>, <string-name><surname>Petrov</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Messer</surname> <given-names>PW</given-names></string-name>. <article-title>Heterozygote advantage as a natural consequence of adaptation in diploids</article-title>. <source>Proceedings of the National Academy of Sciences</source>. <year>2011</year>.</mixed-citation></ref>
<ref id="c53"><label>53.</label><mixed-citation publication-type="journal"><string-name><surname>Karolchik</surname> <given-names>D</given-names></string-name>, <string-name><surname>Hinrichs</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Furey</surname> <given-names>TS</given-names></string-name>, <string-name><surname>Roskin</surname> <given-names>KM</given-names></string-name>, <string-name><surname>Sugnet</surname> <given-names>CW</given-names></string-name>, <string-name><surname>Haussler</surname> <given-names>D</given-names></string-name>, <etal>et al.</etal> <article-title>The UCSC Table Browser data retrieval tool</article-title>. <source>Nucleic Acids Res</source>. <year>2004</year>;<volume>32</volume>(suppl_1):D493-D6.</mixed-citation></ref>
<ref id="c54"><label>54.</label><mixed-citation publication-type="journal"><string-name><surname>Pruitt</surname> <given-names>KD</given-names></string-name>, <string-name><surname>Tatusova</surname> <given-names>T</given-names></string-name>, <string-name><surname>Maglott</surname> <given-names>DR</given-names></string-name>. <article-title>NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins</article-title>. <source>Nucleic Acids Res</source>. <year>2005</year>;<volume>33</volume>(suppl_1):D501-D4.</mixed-citation></ref>
<ref id="c55"><label>55.</label><mixed-citation publication-type="journal"><string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tamura</surname> <given-names>K</given-names></string-name>, <string-name><surname>Sanderford</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gray</surname> <given-names>VE</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>S</given-names></string-name>. <article-title>A Molecular Evolutionary Reference for the Human Variome</article-title>. <source>Mol Biol Evol</source>. <year>2015</year>;<volume>33</volume>(<issue>1</issue>):<fpage>245</fpage>–<lpage>54</lpage>.</mixed-citation></ref>
<ref id="c56"><label>56.</label><mixed-citation publication-type="journal"><string-name><surname>Consortium</surname> <given-names>IH</given-names></string-name>. <article-title>A second generation human haplotype map of over 3.1 million SNPs</article-title>. <source>Nature</source>. <year>2007</year>;<fpage>449</fpage>(7164):851.</mixed-citation></ref>
<ref id="c57"><label>57.</label><mixed-citation publication-type="journal"><string-name><surname>Howe</surname> <given-names>KL</given-names></string-name>, <string-name><surname>Achuthan</surname> <given-names>P</given-names></string-name>, <string-name><surname>Allen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Allen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Alvarez-Jarreta</surname> <given-names>J</given-names></string-name>, <string-name><surname>Amode</surname> <given-names>MR</given-names></string-name>, <etal>et al.</etal> Ensembl <year>2021</year>. <source>Nucleic Acids Res</source>. <volume>2021</volume>;<fpage>49</fpage>(D1):D884-D91.</mixed-citation></ref>
<ref id="c58"><label>58.</label><mixed-citation publication-type="journal"><string-name><surname>Kozlov</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Darriba</surname> <given-names>D</given-names></string-name>, <string-name><surname>Flouri</surname> <given-names>T</given-names></string-name>, <string-name><surname>Morel</surname> <given-names>B</given-names></string-name>, <string-name><surname>Stamatakis</surname> <given-names>A</given-names></string-name>. <article-title>RAxML-NG: a fast, scalable and user- friendly tool for maximum likelihood phylogenetic inference</article-title>. <source>Bioinformatics</source>. <year>2019</year>;<volume>35</volume>(<issue>21</issue>):<fpage>4453</fpage>–<lpage>5</lpage>.</mixed-citation></ref>
<ref id="c59"><label>59.</label><mixed-citation publication-type="journal"><string-name><surname>Cunningham</surname> <given-names>F</given-names></string-name>, <string-name><surname>Allen</surname> <given-names>JE</given-names></string-name>, <string-name><surname>Allen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Alvarez-Jarreta</surname> <given-names>J</given-names></string-name>, <string-name><surname>Amode M</surname> <given-names>R</given-names></string-name>, <string-name><surname>Armean Irina</surname> <given-names>M</given-names></string-name>, <etal>et al.</etal> Ensembl <year>2022</year>. <source>Nucleic Acids Res</source>. <volume>2021</volume>;<fpage>50</fpage>(D1):D988-D95.</mixed-citation></ref>
<ref id="c60"><label>60.</label><mixed-citation publication-type="journal"><string-name><surname>Cingolani</surname> <given-names>P</given-names></string-name>, <string-name><surname>Platts</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>LL</given-names></string-name>, <string-name><surname>Coon</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <etal>et al.</etal> <article-title>A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3</article-title>. <source>Fly</source>. <year>2012</year>;<volume>6</volume>(<issue>2</issue>):80-92.</mixed-citation></ref>
<ref id="c61"><label>61.</label><mixed-citation publication-type="journal"><string-name><surname>Moore</surname> <given-names>JE</given-names></string-name>, <string-name><surname>Purcaro</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Pratt</surname> <given-names>HE</given-names></string-name>, <string-name><surname>Epstein</surname> <given-names>CB</given-names></string-name>, <string-name><surname>Shoresh</surname> <given-names>N</given-names></string-name>, <string-name><surname>Adrian</surname> <given-names>J</given-names></string-name>, <etal>et al.</etal> <article-title>Expanded encyclopaedias of DNA elements in the human and mouse genomes</article-title>. <source>Nature</source>. <year>2020</year>;<volume>583</volume>(7818):<fpage>699</fpage>-710.</mixed-citation></ref>
<ref id="c62"><label>62.</label><mixed-citation publication-type="journal"><string-name><surname>Farrell</surname> <given-names>CM</given-names></string-name>, <string-name><surname>Goldfarb</surname> <given-names>T</given-names></string-name>, <string-name><surname>Rangwala</surname> <given-names>SH</given-names></string-name>, <string-name><surname>Astashyn</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ermolaeva</surname> <given-names>OD</given-names></string-name>, <string-name><surname>Hem</surname> <given-names>V</given-names></string-name>, <etal>et al.</etal> <article-title>RefSeq Functional Elements as experimentally assayed nongenic reference standards and functional interactions in human and mouse</article-title>. <source>Genome Research</source>. <year>2022</year>;<volume>32</volume>(<issue>1</issue>):<fpage>175</fpage>–<lpage>88</lpage>.</mixed-citation></ref>
<ref id="c63"><label>63.</label><mixed-citation publication-type="journal"><string-name><surname>Lesurf</surname> <given-names>R</given-names></string-name>, <string-name><surname>Cotto</surname> <given-names>KC</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Griffith</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kasaian</surname> <given-names>K</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>SJ</given-names></string-name>, <etal>et al.</etal> <article-title>ORegAnno 3.0: a community- driven resource for curated regulatory annotation</article-title>. <source>Nucleic Acids Res</source>. <year>2016</year>;<volume>44</volume>(<issue>D1</issue>):<fpage>D126</fpage>–<lpage>D32</lpage>.</mixed-citation></ref>
<ref id="c64"><label>64.</label><mixed-citation publication-type="journal"><string-name><surname>Box</surname> <given-names>GE</given-names></string-name>, <string-name><surname>Cox</surname> <given-names>DR</given-names></string-name>. <article-title>An analysis of transformations</article-title>. <source>Journal of the Royal Statistical Society Series B: Statistical Methodology</source>. <year>1964</year>;<volume>26</volume>(<issue>2</issue>):<fpage>211</fpage>–<lpage>43</lpage>.</mixed-citation></ref>
<ref id="c65"><label>65.</label><mixed-citation publication-type="journal"><string-name><surname>Spence</surname> <given-names>JP</given-names></string-name>, <string-name><surname>Song</surname> <given-names>YS</given-names></string-name>. <article-title>Inference and analysis of population-specific fine-scale recombination maps across 26 diverse human populations</article-title>. <source>Science Advances</source>. <year>2019</year>;<volume>5</volume>(<issue>10</issue>):eaaw9206.</mixed-citation></ref>
<ref id="c66"><label>66.</label><mixed-citation publication-type="journal"><string-name><surname>Wright</surname> <given-names>S</given-names></string-name>. <article-title>Coefficients of inbreeding and relationship</article-title>. <source>Amer Nat</source>. <year>1922</year>;<volume>56</volume>:<fpage>330</fpage>–<lpage>8</lpage>.</mixed-citation></ref>
<ref id="c67"><label>67.</label><mixed-citation publication-type="journal"><string-name><surname>Kelleher</surname> <given-names>J</given-names></string-name>, <string-name><surname>Thornton</surname> <given-names>KR</given-names></string-name>, <string-name><surname>Ashander</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ralph</surname> <given-names>PL</given-names></string-name>. <article-title>Efficient pedigree recording for fast population genetics simulation</article-title>. <source>PLOS Computational Biology</source>. <year>2018</year>;<volume>14</volume>(<issue>11</issue>):<fpage>e1006581</fpage>.</mixed-citation></ref>
<ref id="c68"><label>68.</label><mixed-citation publication-type="journal"><string-name><surname>Adrion</surname> <given-names>JR</given-names></string-name>, <string-name><surname>Cole</surname> <given-names>CB</given-names></string-name>, <string-name><surname>Dukler</surname> <given-names>N</given-names></string-name>, <string-name><surname>Galloway</surname> <given-names>JG</given-names></string-name>, <string-name><surname>Gladstein</surname> <given-names>AL</given-names></string-name>, <string-name><surname>Gower</surname> <given-names>G</given-names></string-name>, <etal>et al.</etal> <article-title>A community- maintained standard library of population genetic models</article-title>. <source>Elife</source>. <year>2020</year>;<volume>9</volume>.</mixed-citation></ref>
<ref id="c69"><label>69.</label><mixed-citation publication-type="journal"><string-name><surname>Tennessen</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Bigham</surname> <given-names>AW</given-names></string-name>, <string-name><surname>O’Connor</surname> <given-names>TD</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Kenny</surname> <given-names>EE</given-names></string-name>, <string-name><surname>Gravel</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal> <article-title>Evolution and functional impact of rare coding variation from deep sequencing of human exomes</article-title>. <source>Science</source>. <year>2012</year>;<volume>337</volume>(6090):<fpage>64</fpage>-9.</mixed-citation></ref>
<ref id="c70"><label>70.</label><mixed-citation publication-type="journal"><string-name><surname>Mi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Thomas</surname> <given-names>P</given-names></string-name>. <article-title>PANTHER pathway: an ontology-based pathway database coupled with data analysis tools</article-title>. <source>Protein networks and pathway analysis: Springer</source>; <year>2009</year>. p. <fpage>123</fpage>–<lpage>40</lpage>.</mixed-citation></ref>
<ref id="c71"><label>71.</label><mixed-citation publication-type="journal"><string-name><surname>Thomas</surname> <given-names>PD</given-names></string-name>, <string-name><surname>Ebert</surname> <given-names>D</given-names></string-name>, <string-name><surname>Muruganujan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mushayahama</surname> <given-names>T</given-names></string-name>, <string-name><surname>Albou</surname> <given-names>LP</given-names></string-name>, <string-name><surname>Mi</surname> <given-names>H</given-names></string-name>. <article-title>PANTHER: Making genome-scale phylogenetics accessible to all</article-title>. <source>Protein Science</source>. <year>2022</year>;<volume>31</volume>(<issue>1</issue>):<fpage>8</fpage>–<lpage>22</lpage>.</mixed-citation></ref>
<ref id="c72"><label>72.</label><mixed-citation publication-type="journal"><string-name><surname>Mi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Muruganujan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ebert</surname> <given-names>D</given-names></string-name>, <string-name><surname>Mills</surname> <given-names>C</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>X</given-names></string-name>, <etal>et al.</etal> <article-title>Protocol Update for large-scale genome and gene function analysis with the PANTHER classification system (v. 14.0)</article-title>. <source>Nature protocols</source>. <year>2019</year>;<volume>14</volume>(<issue>3</issue>):703-21.</mixed-citation></ref>
<ref id="c73"><label>73.</label><mixed-citation publication-type="journal"><string-name><surname>Mafessoni</surname> <given-names>F</given-names></string-name>, <string-name><surname>Grote</surname> <given-names>S</given-names></string-name>, <string-name><surname>de Filippo</surname> <given-names>C</given-names></string-name>, <string-name><surname>Slon</surname> <given-names>V</given-names></string-name>, <string-name><surname>Kolobova</surname> <given-names>KA</given-names></string-name>, <string-name><surname>Viola</surname> <given-names>B</given-names></string-name>, <etal>et al.</etal> <article-title>A high-coverage Neandertal genome from Chagyrskaya Cave</article-title>. <source>Proceedings of the National Academy of Sciences</source>. <year>2020</year>;<volume>117</volume>(<issue>26</issue>):<fpage>15132</fpage>–<lpage>6</lpage>.</mixed-citation></ref>
<ref id="c74"><label>74.</label><mixed-citation publication-type="journal"><string-name><surname>Prüfer</surname> <given-names>K</given-names></string-name>, <string-name><surname>de Filippo</surname> <given-names>C</given-names></string-name>, <string-name><surname>Grote</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mafessoni</surname> <given-names>F</given-names></string-name>, <string-name><surname>Korlević</surname> <given-names>P</given-names></string-name>, <string-name><surname>Hajdinjak</surname> <given-names>M</given-names></string-name>, <etal>et al.</etal> <article-title>A high-coverage Neandertal genome from Vindija Cave in Croatia</article-title>. <source>Science</source>. <year>2017</year>;<volume>358</volume>(6363):<fpage>655</fpage>-8.</mixed-citation></ref>
<ref id="c75"><label>75.</label><mixed-citation publication-type="journal"><string-name><surname>Prufer</surname> <given-names>K</given-names></string-name>, <string-name><surname>Racimo</surname> <given-names>F</given-names></string-name>, <string-name><surname>Patterson</surname> <given-names>N</given-names></string-name>, <string-name><surname>Jay</surname> <given-names>F</given-names></string-name>, <string-name><surname>Sankararaman</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sawyer</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal> <article-title>The complete genome sequence of a Neanderthal from the Altai Mountains</article-title>. <source>Nature</source>. <year>2014</year>;<volume>505</volume>(7481):<fpage>43</fpage>-9.</mixed-citation></ref>
<ref id="c76"><label>76.</label><mixed-citation publication-type="journal"><string-name><surname>Meyer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kircher</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gansauge</surname> <given-names>M-T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Racimo</surname> <given-names>F</given-names></string-name>, <string-name><surname>Mallick</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal> <article-title>A high-coverage genome sequence from an archaic Denisovan individual</article-title>. <source>Science</source>. <year>2012</year>;<volume>338</volume>(6104):<fpage>222</fpage>-6.</mixed-citation></ref>
<ref id="c77"><label>77.</label><mixed-citation publication-type="journal"><string-name><surname>Browning</surname> <given-names>SR</given-names></string-name>, <string-name><surname>Browning</surname> <given-names>BL</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tucci</surname> <given-names>S</given-names></string-name>, <string-name><surname>Akey</surname> <given-names>JM</given-names></string-name>. <article-title>Analysis of human sequence data reveals two pulses of archaic Denisovan admixture</article-title>. <source>Cell</source>. <year>2018</year>;<volume>173</volume>(<issue>1</issue>):<fpage>53</fpage>–<lpage>61</lpage>. e9.</mixed-citation></ref>
</ref-list>
<sec id="d1e3760">
<title>Supplementary Figures &amp; Tables</title>
<p>Supplementary Table 1. Results of simulation-based tests of dispersion of positive ΔEP SNPs.</p>
<p>Supplemental Table 2. Gene ontology results</p>
<p>Supplemental Table 3. FST Values across ΔEP spectrum of values. Mean FST rank value for UK10K SNPs in ΔEP bins for three population contrasts. Values are for SNPs that are in the UK10K sample and occur with at least 10 derived alleles in the pooled populations of the contrast. For each ΔEP SNP the observed FST was ranked against that for control alleles of the same derived allele frequency.</p>
<p>Supplementary Table 4. ΔEP measures for fixed and polymorphic alleles. Based on maximum-likelihood rooting estimates of ancestral alleles (see <xref rid="fig1" ref-type="fig">Figure 1</xref> for values based on Ensembl rooting). Simulated mean ΔEP was calculated for each SNP by considering all possible non- synonymous mutations and the corresponding EP value for the resulting amino acid in proportion to their mutation probabilities based on empirical estimates. 95% confidence intervals on the mean, determined by bias-corrected bootstrap, are given in parentheses.</p>
<p>Supplementary Table 5. Statistical power for detecting excess heterozygosity.</p>
<p>Supplementary Figure 1. Distributions of derived polymorphism frequency in UK10K.</p>
<p>Distribution of derived allele frequency for each ΔEP bin from –1 to +1 in 0.1 increments. Derived allele frequency ranges from singletons (1 copy of the derived allele) to 7241 copies (only one copy of the ancestral allele). The majority of sites are found at low frequencies across all bins</p>
</sec>
</back>
<sub-article id="sa0" article-type="editor-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.93258.1.sa3</article-id>
<title-group>
<article-title>eLife Assessment</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Ross-Ibarra</surname>
<given-names>Jeffrey</given-names>
</name>
<role specific-use="editor">Reviewing Editor</role>
<aff>
<institution-wrap>
<institution>University of California, Davis</institution>
</institution-wrap>
<city>Davis</city>
<country>United States of America</country>
</aff>
</contrib>
</contrib-group>
<kwd-group kwd-group-type="evidence-strength">
<kwd>Incomplete</kwd>
</kwd-group>
<kwd-group kwd-group-type="claim-importance">
<kwd>Valuable</kwd>
</kwd-group>
</front-stub>
<body>
<p>Drawing on a human population genomic data set, this <bold>valuable</bold> study seeks to show that potentially advantageous alleles are on average older than neutral alleles, invoking the action of balancing selection as the underlying explanation. Currently it is unfortunately unclear how robust the estimates of allele ages are, and the evidence for the authors' proposal is therefore at this stage <bold>incomplete</bold>. If confirmed, the conclusions would be of interest to population genomicists, especially those studying humans.</p>
</body>
</sub-article>
<sub-article id="sa1" article-type="referee-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.93258.1.sa2</article-id>
<title-group>
<article-title>Reviewer #1 (Public Review):</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<anonymous/>
<role specific-use="referee">Reviewer</role>
</contrib>
</contrib-group>
</front-stub>
<body>
<p>Summary:</p>
<p>
In this study, the authors attempt to reinvestigate an old question in population genetics regarding the age of alleles that have experienced different strengths (and directions) of natural selection. Under simple population genetic models, alleles that are positively selected are expected to change frequency in populations faster than neutral alleles. So the naïve expectation is that if you look at alleles that are the same population frequency, those that have been evolving neutrally should have been segregating in the population longer than those that have been experiencing natural selection. While this is exactly what the authors find for alleles inferred to be experiencing negative selection (i.e. they tend to be younger than alleles inferred to be neutral that are at the same frequency), the authors find the opposite for alleles inferred to be under positive selection: they tend to be older than alleles inferred to be neutral. The authors argue that this pattern can be explained by a model where positively selected mutations experience a phase of balancing selection that can dramatically extend the period of time that these alleles segregate in the population.</p>
<p>Strengths:</p>
<p>
The question that the authors address is very interesting and thought provoking. When confronted with a counter-intuitive finding, the authors describe an interesting hypothesis to explain it. The authors investigate a number of interesting sub analyses to corroborate their findings.</p>
<p>Weaknesses:</p>
<p>
While there are some intriguing hypotheses in this manuscript, I struggle to be convinced. The main point that the authors argue is that positively selected alleles are older than their neutral counterparts at the same frequency. They argue that this may be because the positively selected alleles are stuck in some form of balancing selection for a long time before they switch to a more classical form of directional selection. The form of balancing selection they argue is one caused by linkage to deleterious alleles, which takes time for the beneficial alleles to recombine onto a more neutral background. I would really like to see some simulations that demonstrate this can actually occur on average. Reading this paper brought back memories of the classic Birky and Walsh (1988; PMCID: PMC281982) paper that argued that linkage amongst selected alleles does not impact the substitution rate of linked neutral alleles, but does reduce the substitution rate among beneficial alleles. Their simple simulations in 1988 illuminated how this works, and they developed a simple mathematical model that helped us understand how it works. In the current paper, it seems the authors are arguing for a similar effect, but rather than focus on beneficial alleles that fix, they are focusing on beneficial alleles that are still segregating. These seem like similar stories, but without simulations or a mathematical model, I struggle to gain any insight into why the observation is the way it is (and not simply due to a number of possible confounding effects noted below).</p>
<p>
There are a number of elements to the methods and interpretation that could use clarification.</p>
<p>
• Genetic data. One of the biggest weaknesses of this analysis is the choice of genetic data. The authors use the UK10k dataset, and reference the 2015 paper. Looking at that paper, it seems that the data may be composed of low coverage whole genome sequencing data (7x) and high coverage exome sequence data (80x). It appears that these data were integrated into a single VCF file, similar to the 1000 Genomes Project Phase 3 data. If these are the data that was used, then there are substantial differences between the coding and non-coding variants that are compared. However, it is possible that the authors chose to restrict the analysis to the low coverage WGS data and neglected to indicate it in the methods section. I will assume that this is the case for the rest of the review, but the authors should clarify.</p>
<p>
• Recombination rates. I believe the authors use an LD-based recombination map. While these maps are correlated at the longer physical distances with pedigree maps, there are substantial differences at shorter physical scales. These differences have been argued to be due to the action of natural selection skewing patterns of LD. If that is the case, then some of the observations in this paper are circular. Please confirm similar findings with a pedigree-based recombination map.</p>
<p>
• Recombination rates, pt 2. The authors compare patterns of non-synonymous coding variants to a set of non-coding, non-regulatory SNPs. They argue &quot;these will necessarily have experienced similar mutational and recombinational processes&quot;. I don't know that this is true. There are both distinct recombination patterns and mutational patterns in genes vs non-coding regions of the genome. It would be important to more carefully match coding and non-coding variants based on both recombination as well as the type of nucleotide change. There are substantial differences in CpG composition in coding vs non-coding regions for example. While the authors say &quot;Analyses thought to be sensitive to CpG high mutability were limited to SNPs that did not occur as part of a CpG&quot;, it is quite unclear what where CpGs were included vs excluded.</p>
<p>
• Identifying ancestral vs derived alleles. It is unclear how the authors identified ancestral vs derived alleles (they say &quot;inferred ancestral sequence from Ensembl (1) and a maximum likelihood estimator&quot;. Several studies have shown that ancestral misidentification can cause skews in the site frequency spectrum. If the ancestral state of some fraction of alleles were misidentified, then the estimated allele age would be incorrect. Figure 1B shows that the mean frequency of the alleles with the largest delta-EP tend to be very low. This makes me think that ancestral misidentification may have impacted the results.</p>
<p>
• Figure 2B and C. I do not understand how the median can be so far outside the mean and error bars. The legend does not specify what the error bars are, but I feel the distribution must be shown if it is so skewed that the mean and any definition of error does not include the median.</p>
<p>
• Inferring allele ages. The authors use two methods for estimating allele ages, but focus on GEVA. They use the default parameter of effective population size 10,000. How sensitive is the model to this assumption? It has been shown that different regions of the genome (particularly coding vs neutral non-coding) experience different rates of deleterious mutations, and therefore different rates of background selection. Simple models of background selection would suggest that these regions will therefore have different effective population sizes.</p>
<p>
• Fst analysis. The authors look at Fst among 3 populations as a function of delta-EP compared to frequency-matched control SNPs. They find there is no statistical support for different levels of Fst in any pairwise comparison for any delta-EP bin. It seems strange that alleles with large delta-EP would not show increased Fst compared to control SNPs... If they are indeed positively selected, the assumption must be that they are then positively selected in all populations, which seems unlikely. Alternatively, by considering only narrow allele frequency bins, it is possible that Fst is also being controlled, and therefore this analysis is non-informative. A simulation would help understand what the expected pattern is here.</p>
<p>
• It would be great to show more figures like 2A. You can place the x-axis on a log-scale so that it is easier to view the lower allele frequencies. This plot clearly shows differences among the 3 categories. I am very surprised at the much shorter error bars for negative delta-EP at high frequency compared to positive delta-EP variants... Shouldn't there be very few negative delta-EP alleles at such high frequency?</p>
</body>
</sub-article>
<sub-article id="sa2" article-type="referee-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.93258.1.sa1</article-id>
<title-group>
<article-title>Reviewer #2 (Public Review):</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<anonymous/>
<role specific-use="referee">Reviewer</role>
</contrib>
</contrib-group>
</front-stub>
<body>
<p>The authors provide an analysis showing that the allele ages of putatively advantageous alleles tend to be older than those of neutral alleles. To do this, the authors first classify mutations as either neutral, advantageous or deleterious based on a metric called the 'evolutionary probability' which is correlated to the impact of selection acting on a mutation. Then, the authors quantify the age of the mutations using the GEVA method and they also quantify tc (the time of the ancestral node of the edge carrying the mutation). Interestingly, the authors find that advantageous mutations tend to have an older allele age and an older value of tc compared to neutral mutations. The authors posit some explanations for this result invoking the action of balancing selection.</p>
<p>This is an interesting paper and its results could merit an important change in our conception of how we believe that natural selection is acting on the human genome. I have concerns about some of the analysis presented on this paper that have to do with two main factors: 1) Showing that the estimates of allele ages and tc are robust on the dataset presented (more on this topic here below). 2) Presenting more simulations or analytical theory where the authors can show that the models presented by the authors to explain the results indeed fit the data well. As an example, the authors could perform some simulations (likely using SLiM) under the balancing selection models posited by the authors and then show that they can produce data where the allele ages for deleterious, neutral and advantageous alleles have similar patterns to what is observed on the genomic dataset analyzed.</p>
<p>Major concerns</p>
<p>- What is the impact of multiple mutations on the same site on the estimates of allele ages with GEVA?</p>
<p>- GEVA, which is one of the methods used by the authors, 'overestimates &quot;intermediate&quot; times and underestimates older times' according to Ragsdale and Thornton (2023) MBE. What is the impact of this effect for the analysis performed by the authors? Do RUNTC has any known biases on their estimate of tc?</p>
<p>- Additionally what is the impact of phasing errors on the estimates of allele age presented by the authors?</p>
</body>
</sub-article>
<sub-article id="sa3" article-type="referee-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.93258.1.sa0</article-id>
<title-group>
<article-title>Reviewer #3 (Public Review):</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<anonymous/>
<role specific-use="referee">Reviewer</role>
</contrib>
</contrib-group>
</front-stub>
<body>
<p>In their manuscript, Pivirotto et al. make an unexpected observation that a set of candidate beneficial alleles according to the Evolutionary Probability method (EP) have estimated ages thousands of years older than control alleles of similar frequency and outside of functional segments. To explain this unexpectedly older ages, the authors propose a number of interesting evolutionary processes related to balancing selection, including staggered sweeps.</p>
<p>It is important to first mention that the authors do find that as expected, deleterious alleles are younger than controls. This provides evidence that the allele age estimates used by the authors are of sufficient quality to detect age differences between groups of genes. I am also convinced by the fact that EP can be used to focus on a set of alleles substantially enriched in deleterious ones, given the very clear frequency patterns related to EP.</p>
<p>I have a number of concerns about the manuscript, including one rather serious one.</p>
<p>My main concern is that many of the observations made by the authors could be caused by mispolarization of alleles, where either (i) mostly low frequency derived alleles are mischaracterized as ancestral and the other, actually ancestral allele is mischaracterized as a high frequency derived allele, or (ii) mostly low frequency ancestral alleles are mischaracterized as derived. Unfortunately, the authors do not even mention the risk of mispolarization in their manuscript. This is a serious problem for this manuscript because ancestral alleles annotated as derived are by definition going to generate older age estimates than if they were truly derived. It would be very useful to be able to have a look at the full distribution of allele ages rather than just confidence intervals as in Figure 1. I happen to have experience with mispolarization of high frequency ancestral alleles as derived by a maximum likelihood method, different from the one used by the authors (Keightley et al Genetics 2018), where the mispolarization became visible as a very suspicious SFS with a visible excess of high frequency variants, especially those expected to be functional (because of the relatively larger corresponding supply of low frequency deleterious functional variants). Even if the ML method used by the authors is not the same, mispolarization is still a serious risk. Glémin et al. Genome Research 2015 also found that mispolarization is far from being a negligible issue.</p>
<p>Mispolarization of low frequency alleles may be especially prominent in the case of mispolarized deleterious alleles associated with a very negative delta-EP, that then appear as alleles with a very positive delta-EP. Focusing on high delta-EP alleles may then in fact enrich the dataset in mispolarized alleles that then result in older age estimates. Looking at Figure 1B especially, I am worried by the fact that very high delta-EP values seem to go back to the frequencies observed for very negative delta-EP. This is what mispolarization of low frequency alleles might cause as a pattern, in this case especially low frequency ancestral alleles being misidentified as derived?</p>
<p>The authors can address the possible issue of mispolarization in multiple ways. First, they can use simulations of sequences to estimate amounts of mispolarization based on their polarization approach, using substitutions/mutation rates as realistic as possible.</p>
<p>
Second, the authors could check if there is suspicious symmetry in the distribution of delta-EP between alleles at frequency f and alleles at frequency 1-f. This pattern could be generated by mispolarization.</p>
<p>My second less serious concern has to do with the use of high delta-EP as evidence that alleles are beneficial. The validation set from the Patel &amp; Kumar 2019 paper is arguably small with 24 known selected variants. It does not follow from the fact that a small set of known selected variants have higher delta-EP, that all variants with high delta-EP tend to be beneficial. This is especially true in the case where beneficial variants tend to be rare, and there are then far more variants expected with high delta-EP than there are beneficial variants. I am willing to change my mind on this if the overall results can be shown to be robust after accounting for allele mispolarization.</p>
<p>Third, I like the idea of staggered sweeps to explain the results, but I am wondering if there is any evidence in the literature of interference between deleterious and advantageous variants that the authors could base their proposed explanation on.</p>
<p>Finally, and I realize that it is a bit of a stretch, I am wondering if the authors could better justify their choices of methods to estimate the age of alleles. What about ARGweaver, Relate or tsdate? How do these methods compare with GEVA? From looking at the literature I could not find a direct comparison of the precision of GEVA compared to these other tools, but it may be worth at least discussing that the results could be further put to the test with other available ARG-based tools to estimate allele ages. Wilder Wohns et al. Science 2022 compare the performance of these different ARG methods with ancient DNA data, and in fact find that GEVA does not perform as well as for example Relate or tsdate.</p>
</body>
</sub-article>
</article>