Lecture 9 of the Molecular Genomics series. Measuring inbreeding from the genome, and reading the history in segment lengths.
Inbreeding is treated at length in the Quantitative Genetics series, where it is defined through the pedigree. This lecture takes the genomic view. The two are answers to the same question asked in different ways, and the difference between them is the point.
Inbreeding is the consequence of mating individuals related through one or more common ancestors. In a closed population it is unavoidable: with a finite number of ancestors, every mating eventually involves relatives.
The inbreeding coefficient \(F\) of an individual is the probability that the two alleles it carries at a randomly chosen locus are identical by descent, that is, are copies of the same ancestral allele.
\(F\) ranges from 0, not inbred, to 1, completely inbred with no remaining genetic variation.
The distinction between two ways in which alleles can be "the same" is the conceptual heart of this lecture, and it is where the genomic and pedigree definitions diverge.
All IBD alleles are IBS; the reverse is not true. A genotype file records IBS, because that is all a genotyping assay can see. Inbreeding is defined in terms of IBD. Bridging the gap between what is observed and what is wanted is the whole problem, and the solution is to look not at single loci but at stretches.
Two alleles matching at one SNP is weak evidence of common ancestry. Two chromosomes matching at five hundred consecutive SNPs is not: coincidental matching across a long stretch is vanishingly improbable, so a long homozygous run is strong evidence of descent from a common ancestor. This is the reasoning that turns IBS data into IBD inference.
Inbreeding matters for three distinct reasons, and conflating them causes confusion.
Reduced genetic diversity. Homozygosity rises and heterozygosity declines. This reduces the population's capacity to respond to future selection and to adapt to changing conditions, which is a long-term concern for both commercial and conservation programmes.
Expression of recessive disorders. Deleterious recessive alleles are common but normally masked in heterozygotes. Inbreeding brings them together in homozygous form, so disorders that were effectively absent begin to appear. The relationship is direct: the frequency of affected homozygotes for a recessive allele at frequency \(q\) rises from \(q^{2}\) to \(q^{2}+Fq(1-q)\). For a rare allele this can be a large proportional increase.
Inbreeding depression. A general decline in quantitative traits, most severe for fitness-related traits: fertility, survival, disease resistance, and to a lesser degree production. The decline is roughly linear in \(F\), so it accumulates steadily rather than appearing at a threshold, which is part of why it is easy to overlook until it is substantial.
| Pedigree-based \(F_{\text{PED}}\) | Genomic \(F_{\text{ROH}}\) | |
|---|---|---|
| What it measures | The expected proportion of the genome that is autozygous | The realised proportion actually observed |
| Depends on | Pedigree depth and accuracy | Marker density and detection settings |
| Founders | Assumed unrelated, so inbreeding before the pedigree starts is invisible | No such assumption; ancient and recent autozygosity are both visible |
| Detects | Mainly recent inbreeding | Recent and historical, distinguishable by segment length |
| Requires | Recorded ancestry | Genotypes only |
Two implications follow. In populations without recorded pedigrees, which includes most smallholder livestock in the world, the genomic estimate is the only estimate available. And even where pedigrees are excellent, the genomic estimate is more informative, because it measures what actually happened rather than what was expected on average.
The offspring of a half-sib mating has \(F_{\text{PED}} = 0.125\) by construction. Its realised autozygosity depends on which segments meiosis happened to transmit, and varies substantially around that expectation. Full sibs from such a mating can differ in \(F_{\text{ROH}}\) by more than the pedigree value itself.
Because inbreeding depression depends on realised autozygosity and not on its expectation, this variation is biologically real and not measurement noise.
A run of homozygosity (ROH) is a contiguous stretch of the genome in which the copies inherited from the two parents are identical. They are identical because both parents inherited that segment from a common ancestor at some point in the past.
ROH are tens of thousands to millions of base pairs long. Every individual carries some, because going far enough back everyone is related. Their total length and their length distribution are what distinguish a highly inbred animal from an outbred one.
Since \(F\) is the proportion of the genome that is autozygous, and ROH are the autozygous segments, the genomic inbreeding coefficient is simply the fraction of the genome they cover.
where \(k\) is the number of ROH segments detected in the animal, \(L_{\text{ROH}_i}\) is the length of segment \(i\), and \(L_{\text{genome}}\) is the total length of the genome covered by the marker panel. Both definitions of \(F\) given earlier are captured: the probability that a randomly sampled locus is autozygous, and the proportion of the genome that is autozygous, are the same number.
An animal is genotyped on a panel covering 2,500 Mb of autosome. ROH detection returns segments of 42, 31, 28, 19, 15, 12, 9 and 6 Mb.
An \(F\) of 6.5% is roughly what a half-sib mating would produce, but the pedigree for this animal might well show much less if the ancestry is incompletely recorded. Note also the segment lengths: the longest are 42 and 31 Mb, which by the reasoning in the next section indicate common ancestors only one or two generations back.
\(F_{\text{ROH}}\) is not parameter-free. A minimum segment length must be chosen, commonly 1, 2, 4, 8 or 16 Mb, along with a minimum number of SNPs and a tolerance for occasional heterozygous calls arising from genotyping error. Different thresholds give different values, and values from studies using different settings are not comparable. The threshold should always be reported.
The length distribution of ROH carries information that a single summary number does not.
An autozygous segment survives only if recombination has not cut it since the common ancestor from whom both copies descend. The more generations, the more meioses, the more opportunities for a crossover to fall inside the segment and end it. Therefore
Long ROH indicate recent common ancestors; short ROH indicate distant ones. Approximately, a segment of length \(L\) Mb points to a common ancestor about \(g \approx 50/L\) generations back.
So an ROH of about 10 Mb points to a common ancestor roughly five generations ago, and an ROH of about 1 Mb to one roughly fifty generations ago, far outside any pedigree.
This makes the ROH length distribution a record of the population's inbreeding history rather than only a measure of its current state. A population with many long ROH is inbreeding now, through recent close matings, which is a management problem with a management solution. A population with many short ROH and few long ones went through a bottleneck in the distant past but is currently mating relatively widely. The two situations require different responses and have identical \(F_{\text{ROH}}\) if the totals happen to match.
Two further uses follow from looking across animals rather than within one.
ROH islands. If a particular genomic region falls inside an ROH in far more animals than chance would predict, that region is an ROH island. Because autozygosity is otherwise scattered more or less at random across the genome, a consistent local excess suggests that something in the region was driven to homozygosity by selection. This is the connection to Lecture 8: ROH islands are one of the standard selection signature statistics, though they confound selection with locally low recombination and with shared demographic history.
Homozygosity mapping. For a recessive disorder, all affected animals must be homozygous at the causal locus, and if the disorder arose from a single mutation they will all be homozygous for the same ancestral haplotype around it. Finding the region that is homozygous and identical across all cases and not across controls localises the causal variant, often from a small number of cases. This is one of the most efficient uses of genomic data in animal breeding, and it has identified the causal variants for a number of livestock recessive disorders.
Pedigree against genome. A bull has \(F_{\text{PED}}=0.05\) from a five-generation pedigree but \(F_{\text{ROH}}=0.14\). Give two explanations, and say how the ROH length distribution would let you distinguish them.
Explanation 1: pedigree depth. \(F_{\text{PED}}\) assumes the founders of the recorded pedigree are unrelated. Five generations back, the founders of a closed breed are in fact related, so all inbreeding accumulated before that point is scored as zero. The genomic estimate has no such assumption and captures it.
Explanation 2: realised variation. Even with a complete pedigree, realised autozygosity varies around its expectation. An animal can simply have inherited more overlapping segments than average.
Distinguishing them. Look at the length distribution. If the excess consists of many short segments, one or two Mb, they trace to ancestors dozens of generations back, well outside the recorded pedigree, and explanation 1 accounts for it. If the excess consists of a few long segments, tens of Mb, they trace to ancestors within the pedigree's span, and either the pedigree contains an error, a misrecorded sire is common, or explanation 2 applies.
In practice explanation 1 dominates in most livestock populations, which is why genomic estimates of inbreeding are routinely higher than pedigree estimates.
Reading a length distribution. Two populations both have mean \(F_{\text{ROH}}=0.10\). In population A the ROH are mostly 15 to 40 Mb; in population B they are mostly 1 to 3 Mb. Describe each population's history and state which needs urgent management intervention.
Population A. Using \(g \approx 50/L\), segments of 15 to 40 Mb point to common ancestors roughly one to three generations back. This is active, ongoing close inbreeding: matings between full sibs, half sibs, or parent and offspring, happening now.
Population B. Segments of 1 to 3 Mb point to common ancestors roughly 17 to 50 generations back. This is an old bottleneck or founder event whose consequences persist, but current matings are not between close relatives.
Population A needs intervention, and it is also the situation where intervention works. The inbreeding is being generated by current mating decisions, so mating plans that avoid close relatives will reduce the rate immediately. Population B's autozygosity is historical and cannot be undone by mating policy; only introducing new genetic material would change it, and that may be undesirable if the population is a distinct genetic resource worth conserving as it is.
This is the clearest argument for reporting the ROH length distribution rather than only \(F_{\text{ROH}}\). The single number gives the same value for two populations requiring opposite responses.
Recessive disorders. A deleterious recessive allele has frequency \(q = 0.03\) in a population. Compute the frequency of affected animals when \(F = 0\) and when \(F = 0.10\), and comment on the proportional change.
The frequency of the recessive homozygote in a population with inbreeding coefficient \(F\) is
With \(F = 0\): \(P(aa) = 0.03^{2} = 0.0009\), that is 9 affected animals per 10,000.
With \(F = 0.10\): \(P(aa) = 0.0009 + 0.10\times0.03\times0.97 = 0.0009+0.00291 = 0.00381\), that is 38 per 10,000.
The frequency is more than four times higher. The proportional increase is severe precisely because the allele is rare: the inbreeding term is linear in \(q\) while the baseline is quadratic, so their ratio grows as \(q\) falls. At \(q=0.001\) an \(F\) of 0.10 raises the incidence by a factor of about 100.
This is why rare recessive disorders appear suddenly in populations undergoing inbreeding, having been effectively invisible before, and why they are one of the first practical consequences a breeding organisation notices.
An ROH island. A 4 Mb region on one chromosome is inside an ROH in 68% of animals, against a genome-wide average of 9%. Before concluding it is a selection signature, what else could produce this, and what would you check?
Low local recombination. ROH end where crossovers fall. A region with an unusually low recombination rate will be inside long segments in many animals for purely mechanical reasons. Check against a recombination map for the species; this is the most common alternative explanation and the easiest to test.
Low marker density or poor marker quality. ROH detection algorithms declare a run when they see consecutive homozygous calls. A region with sparse markers, or with many markers failing quality control, produces long apparent runs because there are too few opportunities to observe a heterozygote. Check SNP density and call rate across the window.
Low local diversity from ascertainment. If the chip contains few markers that are polymorphic in this population at this location, homozygosity is guaranteed by the panel design rather than by the biology. Check minor allele frequencies of the markers in the window in this population specifically.
Shared demographic history. A recent bottleneck raises autozygosity everywhere, and by chance some regions more than others. Check whether the 68% is a genuine outlier against the empirical distribution of window-wise ROH frequency in this population, not against an absolute figure.
Only after these does a selection interpretation become reasonable, and it is strengthened considerably if the same region is an ROH island in an independently selected population.