Lecture 4 of the Quantitative Genetics series. Identity by descent, coancestry, the relationship matrix, inbreeding and assortative mating.
Two individuals are related if
The first point is frequently overlooked and is a common source of confusion. All members of a population share ancestors at some point in the phylogeny, so relatedness is undefined without a stated reference point. In practice we consider recent ancestry rather than ancient common ancestry: if no further information is available for the top generation of a pedigree, those individuals are taken as the base animals and treated as unrelated and non-inbred. Every relationship coefficient you compute is therefore a statement about ancestry since that base.
The converse does not hold: a homozygote need not be inbred, since the two identical-looking alleles may descend from different ancestral copies.
Individuals in a population vary considerably. Those differences may arise from variation in genes, from environmental experience, or from both. For most traits it is both. To the extent that a trait is influenced by heritable effects, phenotypic similarity should track genetic similarity. We can make that precise by decomposing the covariance between two individuals' phenotypes.
This result is the basis of quantitative genetic estimation from relatives, and also its principal limitation: relatives usually share environments as well as genes, so \(\operatorname{Cov}(e_i,e_j)\ne 0\) inflates estimated heritability unless the design controls for it.
Take a random allele at a locus from individual \(i\) and a random allele at the same locus from individual \(j\).
Ignoring epistasis, split each genotypic value into its additive and dominance parts, \(g_i=a_i+\delta_i\):
The dominance term \(\operatorname{Cov}(\delta_i,\delta_j)\) is normally only appreciable if \(i\) and \(j\) are full sibs, since only full sibs can share both alleles at a locus IBD. Generally, therefore, we focus on \(\operatorname{Cov}(a_i,a_j)\).
For a few common outbred relative pairs this gives:
| Relationship | \(\theta_{ij}\) | \(\operatorname{Cov}(a_i,a_j)\) |
|---|---|---|
| Parent and offspring | \(1/4\) | \(\tfrac12 V_a\) |
| Half sibs | \(1/8\) | \(\tfrac14 V_a\) |
| Full sibs | \(1/4\) | \(\tfrac12 V_a\) (plus \(\tfrac14 V_d\)) |
Parents have breeding values \(a_s\) and \(a_d\). An offspring's expected deviation from the population mean is the parent average breeding value:
An offspring is not exactly the average of its parents. Which particular allele each parent transmits is a coin toss, and that randomness has a name and a variance. With no selection and random mating, \(V_a\) is constant across generations, so:
Given parents with breeding values \(a_s\) and \(a_d\), the two sources of variance among their offspring are \(V_M=\tfrac12 V_a\) and \(V_e\). Hence
This is why even an elite mating produces a range of offspring, and why selection among full sibs is possible at all.
If the two parents are related, there is a chance they transmit IBD alleles to their offspring. That probability is the inbreeding coefficient \(F_i\) of the offspring. Note the logic: relatedness is a property of the parents, inbreeding is a property of the offspring.
\(F\) can be computed from pedigree data in two ways: by direct computation of IBD probabilities using Wright's path method, or by the tabular method, which leads to the numerator relationship matrix below.
Count the number of meiosis links in each loop through a common ancestor:
and if the ancestor in loop \(i\) is itself inbred with coefficient \(F_{A_i}\),
\(n_i\) is the number of meioses (links) in loop \(i\). Each meiosis halves the probability of transmitting the same allele, which is where the factor of one half comes from.
When an individual is inbred, its two alleles \(A_1\) and \(A_2\) at a locus are no longer independent, and neither are their effects \(\alpha_1\) and \(\alpha_2\).
So inbreeding increases the additive genetic variance among individuals by a factor \((1+F)\). This is not extra genetic variation being created; it is the same variation redistributed, more of it now expressed between individuals because fewer alleles are hidden in heterozygotes.
Inbred parents generate less variation among their offspring, because at loci where a parent is already homozygous by descent, segregation produces nothing new:
Start from a population of non-inbred animals. In the next generation, 10% of matings are between half sibs and 5% between full sibs. Half-sib mating gives offspring \(F=0.125\); full-sib mating gives \(F=0.25\). Then
Let a fraction \(\bar F\) of individuals be inbred at a locus. Among the \(1-\bar F\) that are not, genotypes are in Hardy-Weinberg proportions \(p^2\), \(2pq\), \(q^2\). Among the \(\bar F\) that are, the two alleles are IBD, so the individual is homozygous: a fraction \(p\) are \(AA\) and \(q\) are \(aa\), with no heterozygotes.
A population shows 50% \(A_1A_1\), 20% \(A_1A_2\), 30% \(A_2A_2\).
\(p=0.5+0.2/2=0.6\), so \(q=0.4\). Expected heterozygosity \(H_{\text{exp}}=2(0.6)(0.4)=0.48\), observed \(H_{\text{obs}}=0.20\).
Take \(\bar F=0.025\) from the earlier example, and a recessive disease allele at \(q=0.001\).
Under Hardy-Weinberg the disease frequency would be \(q^2=10^{-6}\). With inbreeding it becomes \(q^2+\bar F pq\approx 2.6\times10^{-5}\).
An average inbreeding of just 2.5% raises the incidence roughly 26-fold. This is why rare recessive defects surface in closed nucleus herds long before inbreeding depression in production traits becomes obvious.
The general observation is that inbred individuals perform worse, with an approximately linear decrease in performance as \(F\) increases.
Formally, inbreeding depression occurs if the dominance deviations \(\delta_{Aa}\), summed across loci, are typically positive, that is, if heterozygotes tend to be above the midpoint of the two homozygotes.
For breeding value estimation we need a matrix representation of all pairwise relationships. The natural one is the numerator relationship matrix \(\mathbf{A}\):
The off-diagonal element \(F_{ij}\) (\(i\ne j\)) is the coefficient of relatedness \(F_{ij}=2\theta_{ij}\); the diagonal element is \(1+F_{ii}\), where \(F_{ii}\) is the individual's inbreeding coefficient. The diagonal exceeding one is exactly the \((1+F)V_a\) result derived above.
\(\mathbf{A}\) is what makes BLUP possible: the additive genetic covariance structure of an entire pedigree is \(\mathbf{A}V_a\). In genomic prediction the same role is played by a genomic relationship matrix \(\mathbf{G}\), estimated from markers rather than deduced from the pedigree.
What happens when we mate the best with the best, phenotypically? This is what breeding programmes usually do, so the consequences matter. Assortative mating means \(\operatorname{Cov}(y_s,y_d)=r>0\).
In the first generation assortative mating increases the additive variance. There is, however, a counteracting effect. Mating superior with superior and inferior with inferior induces inbreeding at the genes affecting the trait, which reduces \(V_M\). Over time the effect reverses:
| Generation | Additive genetic variance | Heritability |
|---|---|---|
| 0 | \(V_a\) | \(h^2\) |
| 1 | \(V_a\left(1+\tfrac12 rh^2\right)\) | \(h^2\dfrac{1+\tfrac12 rh^2}{1+\tfrac12 r\left(h^2\right)^2}\) |
| \(\infty\) | \(V_a\left(1-rh^2\right)\) | \(h^2\dfrac{1-rh^2}{1+r\left(h^2\right)^2}\) |
Regression on parent average. Show that the regression of an offspring's phenotype on the parent average phenotype equals the heritability, that is \(\operatorname{Reg}(y_o,\overline{P})=h^2\), where \(\overline{P}=\tfrac12(y_s+y_d)\).
You may use the result from the lecture that the covariance between one parent and its offspring is \(\tfrac12 V_a\).
We need \(\operatorname{Reg}(y_o,\overline{P})= \dfrac{\operatorname{Cov}(y_o,\overline{P})}{\operatorname{Var}(\overline{P})}\), so compute both.
Denominator. Assuming the parents are unrelated and each has phenotypic variance \(V_p\),
Numerator. Using \(\operatorname{Cov}(y_o,y_s)=\operatorname{Cov}(y_o,y_d)=\tfrac12 V_a\),
Ratio.
Note the contrast with regression on a single parent, which gives \(\tfrac12 h^2\). Averaging the two parents halves the variance of the predictor while leaving the covariance unchanged, which doubles the regression coefficient.
Path method. Two half sibs share a single common ancestor and are mated together. Compute the inbreeding coefficient of their offspring, first assuming the common ancestor is not inbred, then assuming it has \(F_A=0.20\).
Trace the loop from one parent up to the common ancestor and back down to the other parent. For half sibs mated together the loop passes through 3 meioses, so \(n=3\).
With an inbred common ancestor,
The ancestor's own inbreeding raises the chance that the two alleles it passes down are IBD, so it inflates the descendant's \(F\) proportionally.
Heterozygote deficiency. A population is scored at one locus: 36% \(A_1A_1\), 48% \(A_1A_2\), 16% \(A_2A_2\). Is there evidence of inbreeding? Now repeat for a second locus scored as 45% \(A_1A_1\), 30% \(A_1A_2\), 25% \(A_2A_2\).
Locus 1. \(p=0.36+0.48/2=0.60\), \(q=0.40\). \(H_{\text{exp}}=2(0.6)(0.4)=0.48\), and \(H_{\text{obs}}=0.48\). \(\bar F=(0.48-0.48)/0.48=0\). No evidence of inbreeding; the locus is in Hardy-Weinberg proportions.
Locus 2. \(p=0.45+0.30/2=0.60\), \(q=0.40\). \(H_{\text{exp}}=0.48\), \(H_{\text{obs}}=0.30\).
A substantial apparent inbreeding. Note the caution, though: both loci have identical allele frequencies, so the difference is entirely in the genotype distribution. A heterozygote deficiency at a single locus can also be produced by population substructure (the Wahlund effect), null alleles, or genotyping error, so \(\bar F\) estimated this way should be averaged over many loci before it is believed.