Selection signatures

Lecture 8 of the Molecular Genomics series. Finding what selection did, without phenotypes.

Download PDF6 pages · 163 KB

← Knowledge Hub

Lecture 8Molecular GenomicsMSc levelGabor Meszaros
Selection signaturesThe full lecture on the Genomics Boot Camp channel.Watch on YouTube ↗

Where we are

A GWAS needs phenotypes. That is a serious constraint, because phenotypes are the expensive part and because selection that acted in the past has already fixed the alleles it favoured, leaving no variation to associate with anything. Selection signature analysis addresses both problems by looking for the mark selection leaves on the genome itself.

What a selection signature is

Definition

Signatures of selection are regions of the genome that harbour functionally important sequence variants and are, or have been, under natural or artificial selection, leaving characteristic patterns of DNA behind. (after Qanbari and Simianer, 2014)

Every population is under selection, natural or artificial or both. Selection changes allele frequencies at the loci it acts on, and because of linkage it changes them in the surrounding region as well. Those changes alter the pattern of nucleotide variation in ways that persist and can be detected long after the episode of selection ended.

The pattern to look for follows from what selection does. Combinations of alleles at very close markers reflect ancestral haplotypes. How long those haplotypes are, and how conserved they are across the population, depends on recombination rate, mutation rate, population size and selection. The first three set a background expectation; selection produces local departures from it.

Genetic hitchhiking

The mechanism is straightforward and worth stating carefully because everything else follows from it.

The hitchhiking argument
  1. A beneficial mutation arises on one particular chromosome, which already carries a particular combination of alleles at all the surrounding loci: one specific haplotype.
  2. Selection increases the frequency of the beneficial allele. Because the allele is physically attached to its haplotype, the whole haplotype rises in frequency with it.
  3. Recombination gradually separates the beneficial allele from the flanking alleles, but this takes time. If selection is strong the allele reaches high frequency before recombination has done much work.
  4. The result is a region around the selected site in which one haplotype is unusually common, variation is unusually low, and homozygosity is unusually high.
beforeduringafterbeneficial mutation arisesreduced variation remains here
Genetic hitchhiking. The beneficial allele (red) sweeps to high frequency carrying its original haplotype (blue) with it. Recombination erodes the swept region from the edges inward, so the reduction in variation is narrowest at the selected site and fades with distance.

The signal is therefore not the beneficial allele itself, which we usually cannot identify, but the depression in variation it leaves in its neighbourhood. This is why the effect is called hitchhiking: the flanking variants contribute nothing and are carried along.

Hard sweeps, soft sweeps and standing variation

The clean picture above is a hard sweep, and it is the easiest case to detect. It is not the only one, and the alternatives are both more common in livestock and harder to find.

TypeOrigin of the beneficial alleleDetectability
Hard, or classic, sweepA single new mutation on a single haplotype Strong signal: one haplotype at high frequency, sharply reduced variation
Soft sweep from standing variationAn allele already present at low frequency in the population when selection began, therefore already on several haplotype backgroundsWeaker: several haplotypes rise together, so variation is reduced much less
Multiple-origin soft sweepThe same beneficial mutation arising independently more than once, on different backgroundsWeakest: no single haplotype is at high frequency at all

Livestock selection has largely acted on standing variation. Breeders selected on phenotypes that were already varying, which means the favoured alleles were already segregating, not newly mutated. So the dominant mode in domestic animals is the soft sweep, and methods calibrated on hard sweeps will systematically miss much of what happened. This is a real limitation, not a technicality: the absence of a detected signature is weak evidence that a region was not selected.

How this differs from GWAS

The two key advantages
  • Phenotypes are not required. The analysis uses genotypes alone. Where phenotype recording is weak, which describes most livestock populations in most of the world, this is decisive.
  • Fixed alleles are detectable. GWAS requires variation: a locus at which every animal has the same genotype contributes nothing to an association test. Selection signatures detect exactly those loci, because fixation with a surrounding haplotype block is the signal.

The second point deserves emphasis. Traits that were under strong selection in the distant past, including much of what happened during domestication and during the formation of breeds, have left no segregating variation to associate with. They are invisible to GWAS by construction and visible to selection signature analysis in principle.

The corresponding disadvantage is the mirror image. A GWAS tells you which trait the locus affects, because you tested against a specific phenotype. A selection signature tells you only that something in this region was selected. Attaching a trait to it requires further work, and this is where most of the interpretive difficulty lies.

Selection signatures also work with a wider range of data. Microsatellites, SNP chips and whole genome sequence can all be used, provided a genetic map is available; and pooled DNA samples, in which many animals are sequenced together without individual identification, are sufficient for several methods, which reduces cost substantially.

Detection approaches

Methods divide according to whether they compare within a population or between populations.

Within a population. These look for local departures from the genome-wide background. The usual signals are reduced heterozygosity, an excess of long homozygous stretches, an unusual allele frequency spectrum with too many rare variants, and extended haplotype homozygosity, in which a haplotype at high frequency remains homozygous over an unexpectedly long distance. The last exploits the fact that a haplotype that is both common and long is anomalous: common usually means old, and old usually means broken up by recombination.

Between populations. These look for regions where two populations differ far more than the rest of the genome would predict. The workhorse is \(F_{ST}\), the same statistic used for ancestry informative markers in Lecture 7, computed in windows along the genome. Drift affects the whole genome roughly equally, so a region of exceptionally high \(F_{ST}\) suggests that selection pushed the populations apart there specifically. Comparing the extent of linkage disequilibrium between populations serves a similar purpose.

Runs of homozygosity. A separate approach identifies long homozygous segments in individual animals and asks whether particular genomic regions are homozygous in far more animals than expected. Such regions are called ROH islands, and they are a hybrid signal: driven partly by selection and partly by inbreeding. This connects directly to Lecture 9.

In practice results are presented much as a GWAS is, as a statistic plotted along the genome with peaks marking candidate regions. The interpretation is analogous and so are the difficulties.

Interpreting the result

Selection signature studies are easy to run and hard to interpret. Three cautions are worth carrying into any reading of the literature.

Demography mimics selection. Bottlenecks, founder events and admixture all reduce variation and extend haplotypes, and they do so across the genome. Distinguishing a selective sweep from a demographic event requires comparison against a background expectation that accounts for the population's history, and that background is itself estimated with uncertainty.

Recombination rate variation mimics selection. Regions of low recombination have low variation and long haplotypes for reasons that have nothing to do with selection. Lecture 3 established that local rates vary by more than an order of magnitude. A signature detected without reference to the local recombination rate is not trustworthy.

A region is not a gene, and a gene is not a trait. Candidate regions typically contain many genes. The common practice of noting that a region contains a gene with a plausible biological connection to a trait of interest is hypothesis generation, not evidence, and the plausibility of such connections is easy to overestimate after the fact.

Selection signature analysis is at its most convincing when several independent methods converge on the same region, when the same region is found in independent populations under similar selection pressure, and when there is functional evidence beyond proximity to a gene with a suggestive name.

References for this section

Exercises

Exercise 8.1

Sweep width. A beneficial allele goes from rare to fixed in 40 generations under strong selection. Using the rule of thumb from Lecture 3, estimate roughly how wide the region of reduced variation around it will be.

Show solution ▾

A flanking variant remains attached to the beneficial allele only if no recombination separates them during the sweep. Over \(g\) generations, a locus at genetic distance \(c\) Morgans from the selected site is separated with probability approximately \(1-(1-c)^{g} \approx gc\) for small \(c\).

Setting \(gc \approx 1\) marks the distance at which separation becomes likely: \(c \approx 1/40 = 0.025\) M = 2.5 cM. Using 1 cM per Mb, that is about 2.5 Mb on each side, so a swept region of roughly 5 Mb.

The general relationship is that sweep width scales as \(s/g\) where \(s\) relates to selection strength: fast sweeps leave wide signatures, slow ones leave narrow ones. Recent, strong selection is therefore easiest to detect, and also worst localised. Ancient selection leaves a narrow, precise, but faint signal. This trade-off between detectability and resolution is intrinsic.

Exercise 8.2

Hard or soft. A region shows moderately reduced heterozygosity but four distinct haplotypes are each at frequencies around 20%, rather than one at 80%. Is this consistent with selection, and what kind?

Show solution ▾

Yes, and it is characteristic of a soft sweep. Under a hard sweep a single haplotype carrying the new mutation rises, so one haplotype should dominate. Four haplotypes at similar intermediate frequency, together with reduced heterozygosity, points instead to selection acting on an allele that was already segregating on several backgrounds when selection began.

The two soft-sweep mechanisms are indistinguishable from this evidence alone: the allele may have been standing variation present on four haplotypes, or the same mutation may have arisen independently more than once. Distinguishing them requires sequence data to determine whether the causal allele is identical by descent across the four backgrounds.

The practical warning is that a method designed to detect one dominant haplotype will report no signal here. Since livestock selection acted overwhelmingly on standing variation, methods tuned to hard sweeps will miss most of the real selection history of domestic animals.

Exercise 8.3

Ruling out the alternatives. A study reports a strong selection signature on bovine chromosome 6 in a dairy breed, based on reduced heterozygosity in a 3 Mb window. List, in order of importance, the checks you would want before believing it, and say what each would rule out.

Show solution ▾

1. Local recombination rate. Compare against a recombination map for the species. Low-recombination regions have low variation for purely mechanical reasons. This is the single most common false-positive mechanism and the cheapest to check.

2. Genome-wide background. Is the window genuinely an outlier relative to the empirical distribution of the statistic across all windows in this population, or merely low in absolute terms? A population that has been through a bottleneck has low variation everywhere, and only relative outliers are informative.

3. Independent methods. Does an \(F_{ST}\) comparison against a related breed, or an extended haplotype homozygosity test, flag the same window? Different statistics respond to different aspects of the sweep, so convergence is meaningful.

4. Independent populations. Is the same region flagged in other dairy breeds selected for similar goals? Convergent signals across independently selected populations are strong evidence, since demographic accidents are unlikely to coincide.

5. Functional evidence. Only after the above does it become useful to ask what genes lie in the window. Chromosome 6 in cattle contains regions associated with milk traits, so a plausible candidate will be found whether or not the signal is real, which is exactly why this check comes last.

Exercise 8.4

Choosing the method. You have SNP genotypes on 300 animals of an indigenous East African breed and 300 of an imported exotic breed, no phenotypes, and a question about adaptation to heat and disease. Which approach would you use, and what is the main threat to the interpretation?

Show solution ▾

Use a between-population approach, primarily windowed \(F_{ST}\), supported by a comparison of the extent of linkage disequilibrium between the two populations. The design is well suited to it: two populations that have experienced very different selection pressures, no phenotypes needed, and sample sizes adequate for allele frequency estimation.

Within-population methods such as extended haplotype homozygosity can be run alongside on the indigenous population, and agreement between the two strengthens any conclusion.

The main threat is that the two populations differ across their whole genomes because they are different populations with long separate histories, not because of adaptation to heat and disease. Elevated \(F_{ST}\) is expected everywhere; only relative outliers mean anything, and the null distribution of \(F_{ST}\) between two long-separated populations is wide, so extreme values arise by drift alone.

A second threat, specific to this comparison, is ascertainment. The chip's markers were discovered in exotic breeds, so allele frequency estimates in the indigenous population are biased, and \(F_{ST}\) computed from them is not an unbiased estimate of the genome-wide value. Sequence data avoids this and would materially strengthen the study.

Adapted by ASAP-Bio from the Introduction to Genomics slide series by Gabor Meszaros (Genomics Boot Camp), released under GPL-3.0. The sequence and worked examples follow his; the explanatory text was written for this platform.
Ministry of Foreign Affairs of Denmark Danida Fellowship Centre
The project is funded by the Ministry of Foreign Affairs of Denmark and managed by Danida Fellowship Centre.
DANIDA Knowledge and Innovation Programme (KIP) 2025.