Genomic admixture

Lecture 7 of the Molecular Genomics series. Estimating where an animal's ancestry comes from, and doing it cheaply.

Download PDF6 pages · 209 KB

← Knowledge Hub

Lecture 7Molecular GenomicsMSc levelGabor Meszaros
Genomic admixtureThe full lecture on the Genomics Boot Camp channel.Watch on YouTube ↗

Where we are

Population structure was a nuisance in the previous lecture, something to be corrected away. Here it becomes the object of study. The same allele frequency differences between populations that confound an association study are exactly the signal used to work out where an animal's ancestry comes from.

What admixture is

Definition

Genetic admixture occurs when individuals from two or more previously separated populations begin interbreeding, introducing new genetic lineages into each.

Admixture arises in two ways. It happens naturally when a geographic barrier separating two populations disappears and they come into contact. And it happens anthropogenically, which in livestock is the usual case: deliberate crossbreeding to introduce production traits, or undirected introgression from imported animals into local populations.

The second is a central question for African animal genetic resources. Indigenous breeds adapted over centuries to local disease pressure, heat and feed availability are widely crossed with imported exotic breeds in pursuit of higher yield. Knowing how much exotic ancestry a population carries, and in which parts of the genome, is a prerequisite for managing that process rather than merely observing it.

The genomic view

Populations that have been separated for a long time accumulate allele frequency differences through drift, and if their environments differ, through selection. No single locus need be fixed for different alleles in the two populations; what matters is that across many loci the frequencies differ systematically.

An admixed animal inherits chromosome segments from ancestors in more than one source population. Its genotype is therefore a mosaic, and the proportion of its genome derived from each source can be estimated from how well its genotypes match the frequency profile of each.

estimated ancestry proportion per animalpure, pop 1pure, pop 2pure, pop 3admixed animalspopulation 1population 2population 3
Each bar is one animal; the coloured fractions are the estimated proportions of its genome derived from each source population. Reference animals of known pure ancestry anchor the estimates.

This kind of analysis has been applied widely to African livestock. Studies of Ethiopian sheep populations, for example, resolve them into a small number of ancestral components whose proportions vary geographically, in a pattern that tracks known migration and trade routes rather than modern administrative boundaries.

Estimating admixture proportions

The estimation problem is to find, for each animal, the proportions \(q_1,\dots,q_K\) summing to one that describe how much of its genome comes from each of \(K\) source populations, and simultaneously the allele frequencies in each source population that make the observed genotypes most probable.

Two things are worth understanding about this, without going into the algorithms.

\(K\) is chosen, not discovered. The number of source populations is supplied by the analyst. Running the analysis at \(K=2\), \(K=3\) and \(K=4\) gives three different, internally consistent descriptions of the same data. Statistical criteria exist for choosing \(K\), but they are guides rather than answers, and the biologically meaningful value is often not the statistically optimal one.

The components are not necessarily breeds. The method finds axes of allele frequency variation. Those axes correspond to real ancestral populations only if such populations exist and are represented in the sample. Supplying reference animals of known pure ancestry is what anchors the components to interpretable labels; without them the components are statistical constructs that require careful interpretation.

Validating against pedigree

The natural test of a genomic admixture estimate is to compare it with a known pedigree.

Swiss Fleckvieh provides an unusually clean case. The breed was established about forty years ago by crossing Simmental with Red Holstein Friesian, and the herdbook records the proportion of each for every animal. The population contains the full range of crosses, from nearly pure Simmental to nearly pure Red Holstein, so pedigree-derived proportions can be compared directly with genomic estimates across the whole spectrum.

The correlation between the two is high, which establishes that the genomic estimate is measuring what it claims to measure. But the two are not identical, and the discrepancy is informative rather than an error.

Why genomic and pedigree admixture differ

Pedigree gives the expected proportion. An F1 crossed back to Simmental has an expected 75% Simmental ancestry. The realised proportion varies around that expectation because which chromosome segments are transmitted is a random outcome of meiosis. Genomic estimates measure the realised proportion, so they capture variation that pedigree cannot see, and full sibs with identical pedigree proportions can differ genomically by several percentage points.

This is the same distinction as between expected and realised relationship, and it is one of the general advantages of genomic over pedigree information.

Ancestry informative markers

Not all markers contribute equally. A marker with the same allele frequency in both source populations carries no information about ancestry at all; a marker close to fixation for different alleles carries a great deal. This suggests selecting a small, highly informative subset.

The standard criterion is the fixation index \(F_{ST}\), the proportion of total genetic diversity that is due to allele frequency differences among populations rather than within them. For two populations with allele frequencies \(p_1\) and \(p_2\) at a marker, and \(\bar p = (p_1+p_2)/2\),

\[F_{ST} = \frac{\bar p(1-\bar p) - \tfrac12\left[p_1(1-p_1)+p_2(1-p_2)\right]}{\bar p(1-\bar p)}\]

Markers with high \(F_{ST}\) between the pure reference populations are called ancestry informative markers. Selecting the top-ranked few hundred or few thousand yields a panel that estimates admixture proportions nearly as well as the full chip.

The empirical result is striking: in the Swiss Fleckvieh work, admixture estimated from one percent of the SNPs, selected on \(F_{ST}\), correlated with pedigree admixture almost as well as estimates from the complete panel. The practical implication is direct. A cheap custom panel of a few hundred markers can do breed composition testing at a fraction of the cost of a full chip, which matters for routine application in settings where a 50K chip per animal is not affordable.

How few markers, how few reference animals

The second cost is the reference set: pure animals of each source population must be genotyped to estimate the ancestral allele frequencies. How many are needed?

Reducing the reference set from 100 animals per pure breed to 50, then 20, then 10, degrades the estimates surprisingly little. Reasonable accuracy is retained even with small reference panels, because allele frequency at a marker with high \(F_{ST}\) is estimated adequately from few individuals when the frequency difference between populations is large.

What this means in practice

Breed composition analysis does not require a large genomic infrastructure. A few hundred well-chosen markers and ten to twenty reference animals per source population are enough for usable estimates. This puts the method within reach of national programmes that cannot fund population-scale genotyping.

The caveat is representativeness rather than number. Ten reference animals that happen to be closely related, or drawn from one herd, will misrepresent the population's allele frequencies. A small reference set must be deliberately spread across the population's structure, which requires knowing something about that structure beforehand.

Cautions

Ascertainment bias returns. Section 2 of Lecture 2 noted that commercial chips contain markers discovered in commercial breeds. For admixture work this is more than a nuisance: the markers that best distinguish an indigenous population from an exotic one are disproportionately those that were never candidates for the chip. \(F_{ST}\) computed from an ascertained panel is not an unbiased estimate of genome-wide \(F_{ST}\).

Reference populations must be genuinely pure. The whole method rests on the assumption that the reference animals represent unadmixed source populations. Where introgression has been going on for decades, as it has in much of East Africa, finding genuinely unadmixed indigenous reference animals may not be possible, and the estimated indigenous component will then be defined relative to an already partly admixed baseline.

Proportion is not the whole story. Two animals with 25% exotic ancestry may carry that ancestry in entirely different parts of the genome. For conservation and for adaptation, where the exotic ancestry sits matters more than how much of it there is. Local ancestry inference, which assigns ancestry along the chromosome rather than as a genome-wide average, addresses this and is the natural next step.

References for this section

Exercises

Exercise 7.1

Computing F-ST. At a marker, allele \(A\) has frequency 0.90 in an exotic breed and 0.20 in an indigenous population. At a second marker the frequencies are 0.45 and 0.55. Compute \(F_{ST}\) for each and say which would be selected as an ancestry informative marker.

Show solution ▾

Marker 1. \(\bar p = (0.90+0.20)/2 = 0.55\), so \(\bar p(1-\bar p)=0.55\times0.45=0.2475\).

Within-population heterozygosity averaged: \(\tfrac12[0.90\times0.10 + 0.20\times0.80] = \tfrac12[0.09+0.16]=0.125\).

\[F_{ST}=\frac{0.2475-0.125}{0.2475}=0.495\]

Marker 2. \(\bar p = 0.50\), so \(\bar p(1-\bar p)=0.25\). Within-population: \(\tfrac12[0.45\times0.55+0.55\times0.45]=0.2475\).

\[F_{ST}=\frac{0.25-0.2475}{0.25}=0.01\]

Marker 1 would be selected; marker 2 carries almost no information about ancestry despite being perfectly polymorphic in both populations. This illustrates the key point: what makes a marker useful for admixture is the difference in frequency, not the level of variation.

Exercise 7.2

Realised versus expected. Two full sibs both have pedigree ancestry of 50% Simmental and 50% Red Holstein. Genomic estimates give 54% and 47% Simmental. Has an error occurred? Roughly how much variation would you expect?

Show solution ▾

No error. Pedigree gives the expected proportion; the realised proportion varies because meiosis samples which segments are transmitted.

For an F1 the parental contributions are exactly 50:50, but from the second generation onward the realised proportion varies. The standard deviation of realised ancestry depends on the number of independently segregating chromosome segments, which is set by the genetic map length. With a genome of about 30 Morgans and \(g\) generations since admixture, the standard deviation of realised ancestry is on the order of a few percent, and it decreases as more generations of recombination fragment the genome into smaller, more numerous segments.

A spread of 54% against 47% is well within that range. The substantive point is that the genomic estimate is the more useful number: if one sib really does carry more Simmental genome than the other, that is a real biological difference that pedigree cannot represent.

Exercise 7.3

Designing a cheap panel. A national programme wants to test whether smallholder cattle carry exotic ancestry, at a cost of no more than a few euros per animal. Outline a design, and state the single assumption that most threatens it.

Show solution ▾

Design. Genotype a modest reference set at high density, perhaps twenty animals of the indigenous population from several widely separated locations and twenty of the exotic breed, using a 50K chip. Compute \(F_{ST}\) for every marker between the two references and select the several hundred with the highest values, checking that they are spread across all chromosomes so that no region dominates. Commission a custom low-density panel of those markers, which is cheap per sample because the assay is small. Genotype field animals on that panel and estimate admixture proportions using the reference allele frequencies already obtained.

The threatening assumption is that the indigenous reference animals are unadmixed. If crossbreeding has been going on for thirty years, animals sampled as indigenous will themselves carry exotic ancestry. The estimated allele frequency difference between the two references is then smaller than the true difference, the selected markers are less informative than they appear, and every field animal's exotic proportion is systematically underestimated because it is measured against a baseline that is already partly exotic.

Mitigations are to sample reference animals from the most remote areas, to use historical or archived samples where they exist, and to state explicitly that estimates are relative to the contemporary indigenous population rather than to an ancestral one.

Exercise 7.4

Choosing K. An admixture analysis of sheep from six Ethiopian regions is run at \(K=2\) through \(K=6\). At \(K=2\) the split is highland against lowland; at \(K=5\) each region has its own component. A colleague asks which is correct. How would you answer?

Show solution ▾

Both are correct descriptions at different resolutions, and the question as posed does not have an answer. Admixture analysis decomposes allele frequency variation into \(K\) axes; increasing \(K\) resolves finer structure, and there is no privileged level at which the decomposition becomes the truth.

What can be said is which is useful for a given purpose. At \(K=2\) the highland-lowland split is likely the dominant axis of variation, reflecting an old separation with an ecological basis, and it is the right level for a question about broad adaptation. At \(K=5\) the components may reflect recent, local drift among small populations rather than deep ancestry, and they are the right level for questions about which flocks exchange animals.

Two practical checks help. Statistical criteria such as cross-validation error indicate the \(K\) the data support, and should be reported. And stability across independent runs matters: components that appear consistently are more credible than those that vary between runs. A result should be presented across a range of \(K\) with the interpretation stated, not as a single value asserted to be the number of ancestral populations.

Adapted by ASAP-Bio from the Introduction to Genomics slide series by Gabor Meszaros (Genomics Boot Camp), released under GPL-3.0. The sequence and worked examples follow his; the explanatory text was written for this platform.
Ministry of Foreign Affairs of Denmark Danida Fellowship Centre
The project is funded by the Ministry of Foreign Affairs of Denmark and managed by Danida Fellowship Centre.
DANIDA Knowledge and Innovation Programme (KIP) 2025.