Lecture 5 of the Quantitative Genetics series. Drift, effective population size, Wright's F-statistics and the balance between gain and diversity.
In deriving Hardy-Weinberg proportions we assumed the population was infinite, which is what guaranteed the constancy of allele frequencies. Real populations are finite. Once we admit that, allele frequencies stop being constant even in the complete absence of selection.
Genetic drift describes the random change in allele frequency due to random sampling of alleles to form the offspring generation from the parent generation.
Where does the randomness come from?
Without selection, allele frequencies do not change on average. The frequency in the next generation is a random variable whose expectation is the current frequency, but whose variance is not zero. That distinction, no expected change but real variance, is the whole of drift.
Take a population of \(N\) diploid individuals. They carry \(2N\) alleles. Each offspring is formed by picking two parents at random with replacement, and picking one allele from each. There is no selection.
The allele frequency in the next generation is then a random variable. If among the parents there are \(n_t\) copies of \(A_1\) out of \(2N\) genes, each allele in the offspring is \(A_1\) independently with probability \(n_t/2N\), so the count follows a binomial distribution:
In a finite population, this random walk has absorbing boundaries.
Selection may preserve polymorphism, for instance under overdominance, but drift alone will not.
The process of genetic drift can be quantified, but only if we assume simplified conditions. The reference is the Fisher-Wright ideal population:
Therefore all individuals have exactly the same chance of contributing, and mating between gametes is random.
Comparing a real population to this idealisation is what gives the effective population size its meaning.
We measure divergence of allele frequencies by \(F_{ST}\), Wright's drift coefficient. It measures the diversification of sub-populations, and equally how far any one population has diverged from its origin. Let us derive how it accumulates.
For a real population, the effective population size is the size of an ideal population that loses variation at the same rate as the real population does.
The fraction of variance lost in one generation is
reading the numerator as the variance lost going from \(t-1\) to \(t\), and the denominator as the variance still left at \(t-1\).
This is the definition of \(N_e\) used throughout. There are other, nearly equivalent, definitions.
The effective size of most populations is typically much smaller than the real number of individuals. Five mechanisms account for most of the gap.
with the upper bound \(N_e<4\min(N_m,N_f)\). This is very important in domestic animals, where a handful of AI sires may serve an enormous number of females, and in some natural populations.
When parents are chosen at random, offspring numbers are approximately Poisson, \(\sigma^2=2\), and \(N_e\approx N\). With selection among parents \(\sigma^2\gg 2\) and \(N_e\ll N\). In many species only a tiny fraction of individuals reproduce at all, oak trees produce huge numbers of seeds of which very few establish.
This is a harmonic mean, so the smallest values dominate. Founder effects and bottlenecks belong in this category, and they are common in both natural and domesticated populations. One severe bottleneck permanently marks a population's \(N_e\), however much it later recovers in census size.
where \(F\) is the deficiency of heterozygotes due to non-random mating. In domestic animals we deliberately mate non-relatives, so \(F<0\) and \(N_e\) is slightly increased. The magnitude is usually small.
Selection reduces the number of reproducing individuals, and typically means selecting related animals, so it reduces \(N_e\). Quantitatively, with selection intensity \(i\),
The asymmetry is important: gain rises linearly with intensity, whereas the loss of variation rises with its square. Increasing selection intensity therefore yields progressively less gain per unit of diversity lost.
Because both raise average \(F\), they are readily confused when a single coefficient is read from a pedigree or a marker panel. The practical consequence is that a heterozygote deficiency caused by non-random mating disappears in one generation of random mating, whereas variation lost to drift is not recoverable.
When we worked with phenotypes we judged the importance of a factor by the variance it explains. The same logic applies to allele frequencies.
Three populations have \(A\) at frequencies \(p_1=0.2\), \(p_2=0.4\), \(p_3=0.6\).
So 11.1% of the variation is between populations.
These are the same quantity seen from two directions, which is why the same \(F_{ST}\) serves both the population-genetics literature on structure and the breeding literature on loss of variation.
Wright united the description of the two effects. An individual may be IBD at a locus because of genetic drift; if it is not, it may still be IBD because of non-random mating.
The probability \(\left(1-F_{IT}\right)\) of not being inbred is the probability of not being inbred due to non-random mating (\(F_{IS}\)) and not due to drift (\(F_{ST}\)).
The three coefficients are interpreted as follows:
Loss of variation, and therefore \(N_e\), is typically measured per generation. That requires a definition of generation time, and it is arguable that per generation is not the most appropriate scale for comparing species or breeding schemes.
Two definitions are in use:
This matters practically: a scheme that halves the generation interval doubles the rate of gain per year, but it also doubles the rate at which inbreeding accumulates per year.
The superiority of hybrids over the average of their parents is due to the cancellation of inbreeding depression caused by drift. Hybrid vigour occurs when hybrids between two populations outperform the average of the parent populations.
This explains why crossing two long-separated inbred lines works so reliably: each line has drifted to fixation for different deleterious recessives, and the cross restores heterozygosity at all of them.
Recall from the previous lecture that the covariance between the additive genetic effects of two individuals is \(\operatorname{Cov}(a_i,a_j)=F_{ij}V_a\). Importantly, this says nothing about the source of \(F_{ij}\), whether drift or mating between relatives. The algebra is identical; only the interpretation and the reversibility differ.
The endpoint is random. Drift has no direction, so two replicate lines from the same base population fix at different means. This is exactly why drift must be taken into account when designing selection experiments, a divergence between selected and control lines is not evidence of response unless it exceeds what drift alone would produce.
The selection intensity \(i\) is the average phenotypic superiority of selected individuals over the population mean, in units of phenotypic standard deviation.
Stronger selection means more genetic gain, roughly \(\Delta G\propto i\). But selection also reduces the number of reproducing individuals and concentrates them among relatives, so \(\Delta F\propto i^2\).
One of the major objectives of modern breeding planning is finding an optimal balance between \(\Delta F\) and \(\Delta G\).
Because loss goes as \(i^2\) while gain goes as \(i\), the ratio \(\Delta G/\Delta F\) falls as selection intensifies. There is no intensity that is simply "best"; the optimum depends on how the programme values long-term variance against short-term gain. This is the question that optimum contribution selection was invented to answer formally.
Estimating \(N_e\) from divergence. A large population is split into five separate parts of equal size. After 10 generations the allele frequencies at a locus are found to be 0.3, 0.4, 0.5, 0.6 and 0.7. Estimate the average \(N_e\) during the first 10 generations, assuming inbreeding in the base population is zero.
You will need \(F_{ST}=\dfrac{\operatorname{Var}(p_i)}{\bar p(1-\bar p)}\) and, after \(T\) generations, \(1-F_{ST_T}=\left(1-\dfrac{1}{2N_e}\right)^{T}\).
Step 1, compute \(F_{ST}\). \(\bar p=(0.3+0.4+0.5+0.6+0.7)/5=0.5\).
\(\operatorname{Var}(p_i)=\overline{p^2}-\bar p^2\). Here \(\overline{p^2}=(0.09+0.16+0.25+0.36+0.49)/5=0.27\), so \(\operatorname{Var}(p_i)=0.27-0.25=0.02\).
Step 2, solve for \(N_e\). With \(T=10\),
Each sub-population therefore behaved as an ideal population of approximately 60 individuals. This is an estimate of the average \(N_e\) over the 10 generations; where size fluctuates, the harmonic mean is dominated by the smallest generations.
Simulating drift and fixation. Use the course R code to simulate
allele-frequency change in a small population. The parameters are tmax (generations,
large enough that most replicates fix), nrep (repetitions), N (population
size) and p0 (starting frequency).
tmax <- 500; nrep <- 200; N <- 30; p0 <- 0.45
fixtime <- rep(NA, nrep)
plot(x=c(0,tmax), y=c(p0,p0), type="l", ylim=c(0,1),
xlab="Generation", ylab="Allele frequency", main=paste("Ne=",N))
lines(x=c(0,tmax), y=c(0,0)); lines(x=c(0,tmax), y=c(1,1))
1. Each replicate performs a random walk until it hits 0 or 1, where it stays (an absorbing boundary). The fixation time is the generation at which polymorphism disappears; the fixation probability is the proportion of replicates in which the tracked allele reaches 1 rather than 0.
2. With \(N=30\) trajectories wander widely and most fix within a few hundred generations. With \(N=300\) the per-generation steps are much smaller (variance in \(p\) is of order \(p q/2N\)) and most replicates remain polymorphic for far longer. Mean time to fixation scales roughly with \(N_e\).
3. Under pure drift the probability that an allele eventually fixes equals its current frequency, so with \(p_0=0.45\) about 45% of replicates should fix for \(A\). This follows from the fact that allele frequency is a martingale under drift: there is no expected change, so the long-run probability of the absorbing state at 1 must equal the starting value.
Comparing effective population sizes. A nucleus herd has 500 breeding females and 10 breeding sires. In a second scenario, the herd has 250 of each. Compute \(N_e\) for both, and comment on which design conserves more variation.
Using \(N_e\approx\dfrac{4N_mN_f}{N_m+N_f}\).
Scenario 1: \(N_e=\dfrac{4\times 10\times 500}{510}=\dfrac{20000}{510}\approx 39\).
Scenario 2: \(N_e=\dfrac{4\times 250\times 250}{500}=\dfrac{250000}{500}=500\).
Both herds contain 510 and 500 animals respectively, almost identical census sizes, yet the effective sizes differ by more than tenfold. The bound \(N_e<4\min(N_m,N_f)=40\) makes the first scenario's ceiling explicit: with only 10 sires, no number of females can rescue \(N_e\).
This is the principal practical implication for animal breeding: intense male selection, which artificial insemination makes possible and which maximises \(\Delta G\), also reduces \(N_e\) most rapidly.