DNA, genes, genomes and markers

Lecture 1 of the Molecular Genomics series. The structure of DNA, what a gene and a genome are, and why molecular markers work.

Download PDF7 pages · 190 KB

← Knowledge Hub

Lecture 1Molecular GenomicsIntroductoryGabor Meszaros
DNA and genetic markersThe full lecture on the Genomics Boot Camp channel.Watch on YouTube ↗

Why an animal scientist needs genomics

Animal breeding worked for a century without any direct observation of DNA. Selection on phenotypes, pedigrees and best linear unbiased prediction transformed dairy cattle, pigs and poultry long before a single genome was sequenced. The statistical machinery of quantitative genetics deliberately treats the genome as a black box: it is enough to know that relatives resemble each other in a way that can be predicted from the pedigree.

What genomics changes is that the box is now open. Instead of inferring how much of the genome two animals share from their pedigree relationship, we can measure it. Instead of waiting for a bull's daughters to milk before knowing his breeding value, we can genotype him at birth. Instead of describing a breed as a name in a herdbook, we can describe it as a set of allele frequencies and ask where it came from.

This course covers the concepts needed to read, and to do, that kind of work. It is deliberately theoretical. The companion Genomics Boot Camp covers the practical side, handling genotype files, running PLINK, quality control. The two are meant to be taken together: the boot camp will teach you which command to type, this course explains what the output means.

Assumed background
  • School-level biology: cells, chromosomes, meiosis.
  • Basic Mendelian genetics: alleles, genotypes, dominance.
  • Elementary statistics: mean, variance, correlation, a regression line.

Everything beyond that is developed as it is needed.

What the structure of DNA had to explain

By the early 1950s it was accepted that DNA carried heredity, but nobody knew how a molecule could do such a thing. The problem was well posed, which is part of why it was solved so quickly. Any proposed structure had to account for three properties simultaneously.

Three requirements
  • Faithful replication. The molecule must be copyable, so that a daughter cell receives the same information as the parent.
  • Informational content. It must be able to encode something. A perfectly regular crystal carries no information; the molecule needs a variable part.
  • Stability with the capacity to change. Mutation must be possible, otherwise there is no evolution, but it must be rare, otherwise the information degrades faster than it can be used.

Three pieces of evidence were on the table.

The building blocks. DNA was known to consist of a phosphate group, a five-carbon sugar (deoxyribose), and one of four nitrogenous bases. The bases fall into two chemical classes: the purines, adenine and guanine, which have a double-ring structure, and the pyrimidines, cytosine and thymine, which have a single ring. One phosphate plus one sugar plus one base is a nucleotide, the repeating unit of the chain.

Chargaff's rule. Erwin Chargaff had measured base composition across species and found a regularity that held everywhere: the total amount of purine equals the total amount of pyrimidine, and more specifically the amount of adenine equals the amount of thymine, and the amount of guanine equals the amount of cytosine.

\[A = T,\qquad G = C,\qquad \text{hence } A+G = T+C\]

This is a strong hint. A one-to-one quantitative correspondence between two specific pairs of molecules suggests those molecules are physically paired.

X-ray diffraction. Rosalind Franklin fired X-rays at DNA fibres and recorded the scattering pattern on photographic film. The angles at which the rays scattered constrain where the atoms can be. Her results indicated a long, thin molecule with a helical, repeating structure and two components running parallel to one another. This was the decisive constraint: it told Watson and Crick they were looking for a structure with two strands.

The double helix

Watson and Crick did not perform a new experiment. They built physical models until one fitted all three constraints at once, and published the result in a one-page paper in Nature in 1953. Watson, Crick and Wilkins received the Nobel Prize for Physiology or Medicine in 1962.

The structure has an alternating phosphate and deoxyribose backbone, with the bases pointing inward. Two such chains run antiparallel, one in the 5' to 3' direction and the other in the 3' to 5' direction, wound around a common axis. The two chains are held together by hydrogen bonds between the bases, and the pairing is specific: adenine pairs with thymine, guanine pairs with cytosine.

Why the structure answers all three requirements
  • Replication follows immediately. If A always pairs with T and G with C, then each strand fully specifies the other. Separate the strands and each acts as a template.
  • Information lies in the sequence of bases along one strand. The backbone is regular, the sequence is not, so an arbitrary message can be written in a four-letter alphabet.
  • Stability comes from the fact that a base pair is held by two or three hydrogen bonds and buried inside the helix, while change remains possible through rare copying errors.

Chargaff's rule is now not a curiosity but a consequence: every A on one strand forces a T on the other, so the totals must be equal.

From nucleotide to gene to genome

The sequence of bases is not uniformly meaningful. A gene is a stretch of DNA with a defined internal organisation. It begins with a regulatory region that controls when and where transcription starts, contains a signal marking where transcription should end, and in between alternates exons, the segments that end up in the mature transcript and are translated into protein, with introns, the segments that are spliced out.

The genome is the complete set of DNA in an organism. Two facts about it matter for everything that follows. First, it is large: a mammalian genome is roughly three thousand million base pairs, distributed over a set of chromosomes. Second, genes occupy a small minority of it. Most of the genome is intergenic, and much of what is not a gene is nonetheless functional in ways that regulate genes.

Both facts are inconvenient for anyone who wants to find the DNA responsible for a trait. There is a great deal of sequence, and no simple rule tells you which parts matter.

The other 'omes'

The suffix -ome denotes the complete collection of something, and -omics the study of it. The vocabulary has proliferated, sometimes past the point of usefulness, but the main terms are worth knowing because they name genuinely different layers of biology.

TermThe complete collection of
GenomeDNA, that is, all the genetic information of an organism
Epigenomechemical modifications to DNA and histone proteins that alter expression without altering sequence
TranscriptomeRNA molecules in a cell or tissue
Proteomeproteins in a cell, tissue or organism
Metabolomesmall-molecule metabolites
Microbiomemicrobes, and their genes, living in or on the organism
Metagenomegenetic material recovered directly from an environmental sample
Phenomephenotypic traits, as influenced by genome and environment

The layers are ordered, roughly, from cause to consequence. The genome is fixed at conception and the same in every cell; the transcriptome, proteome and metabolome differ between tissues and change by the hour. That difference is exactly why the genome is the convenient place to start. A genotype measured once is valid for life.

This course concerns the genome. The other layers appear in the ASAP-Bio multi-omics theme.

Genetics and genomics

The two words are often used interchangeably and should not be. The World Health Organization draws the distinction as follows.

Definitions
  • Genetics is the study of heredity, and of individual genes: how a gene is composed, what it does, how it is transmitted.
  • Genomics is the study of all the genes of an organism together, their interrelationships, and their combined influence on growth and development, together with the techniques that make such study possible.

The practical difference is one of scale, and scale changes the methods. When you study one gene you can reason about it directly. When you study half a million markers simultaneously you cannot, and statistical problems appear that have no analogue in single-gene genetics: multiple testing, correlations between markers, more variables than observations. A large part of this course is about those problems.

Why genomics happened when it did

The intellectual framework for genomics was largely in place well before the data were. The limiting factor was the cost of finding out what an animal's DNA actually says.

PeriodWhat became possible
1860sMendel establishes particulate inheritance
1953Watson and Crick, the structure of DNA
1970sCloning and sequencing, one gene at a time
1980sGenome projects: concentrated efforts to sequence whole genomes and build genomic maps, allowing overall patterns to be seen rather than isolated genes
1990s to 2010sGenotyping and sequencing become fast, automated and cheap; access moves from a handful of large laboratories to essentially anyone

The decisive steps were the arrival of high-density single nucleotide polymorphism (SNP) genotyping and, later, affordable whole genome sequencing. What mattered was not that these were possible in principle but that they became cheap enough to apply to thousands of animals. Genetic architecture is a population-level question, and population-level questions need population-sized samples.

This produced the pattern typical of a technological shift in a science: old questions were answered, new ones were created, and the new questions required methods that did not exist.

Molecular markers

We cannot look at everything at once, and for most purposes we do not need to. The working compromise of genomics is the molecular marker.

Definition

A molecular marker is a segment of DNA with an identifiable physical location on a chromosome, whose inheritance can be followed. It may be a stretch of sequence or a single base pair.

A marker is not usually interesting in itself. Its value is positional. If we know a marker's location, and we can score its genotype cheaply in many animals, then it acts as a signpost for whatever lies near it. When a marker's genotype is associated with a trait, the reasonable conclusion is not that the marker causes the trait but that something close to the marker does.

The logic is the same as looking for a party in an unfamiliar town. You cannot see the house, but you can see the cars parked along the street. Where the cars are dense, the house is near. The markers are the cars; the gene is the house. How far along the street the cars extend, and therefore how precisely they locate the house, is a question about linkage disequilibrium, which is Lecture 5.

geneSNP 123456789markers 5 and 6 flank the gene and will be associated with itchromosome
Markers are scored across the whole chromosome at known positions. The causal gene is not observed, but the markers nearest to it carry information about it.

The single nucleotide polymorphism

Several marker types have been used historically, including microsatellites and restriction fragment length polymorphisms. The SNP has displaced almost all of them.

Definition

A single nucleotide polymorphism is a single base pair position in genomic DNA at which more than one sequence alternative, that is more than one allele, is present in the population.

Two properties explain the dominance of the SNP. It is abundant, occurring roughly every few hundred to few thousand bases in livestock genomes, so markers can be placed densely enough to cover the whole genome. And it is simple: because it is a single position with, in practice, two alleles, it can be scored by an automated assay with a binary readout, which makes hundreds of thousands of simultaneous measurements economically feasible.

The second property has a cost. A SNP marker carries less information per locus than a microsatellite, which may have ten or more alleles. Genomics compensates with volume: many uninformative markers beat a few informative ones when the many are dense enough to blanket the genome. How that trade-off works out in practice is the subject of Lecture 2.

References for this section

Exercises

Exercise 1.1

Chargaff arithmetic. A double-stranded DNA molecule is found to be 22% adenine. What are the percentages of the other three bases? What would you conclude if a sample were reported as 30% A and 25% T?

Show solution ▾

Because A pairs with T and G with C, in double-stranded DNA \(A=T\) and \(G=C\), and the four must sum to 100%.

If \(A=22\%\) then \(T=22\%\), leaving \(100-44=56\%\) to be divided equally between G and C, so \(G=C=28\%\).

A report of 30% A and 25% T is inconsistent with a double-stranded molecule. Three explanations are worth considering, in roughly this order: measurement error; the material is single-stranded, for example a viral genome or an RNA preparation, in which case Chargaff's rule does not apply; or the sample is contaminated with material of a different composition. The rule is a diagnostic as well as a fact.

Exercise 1.2

Why two strands. Franklin's diffraction data indicated two strands. Suppose the evidence had instead pointed to a single strand. Which of the three requirements in this lecture would become difficult to satisfy, and why?

Show solution ▾

Replication is the requirement that fails. Information content and stability are both achievable with one strand: a single chain of four bases can encode an arbitrary message, and covalent bonds along a backbone are stable.

What complementary double-strandedness supplies is a mechanism for copying. Each strand contains the full information needed to reconstruct the other, so separating them yields two templates rather than one original and one blank. Without complementarity there is no obvious physical basis for accurate duplication, which is why Watson and Crick's closing remark, that the pairing they proposed immediately suggests a copying mechanism, was the substance of the paper rather than an aside.

Exercise 1.3

Markers and genes. A study reports that a SNP is significantly associated with milk yield. A colleague concludes that the SNP causes the difference in yield. Give two reasons why that conclusion is probably wrong, and state what the result does support.

Show solution ▾

First, SNPs on commercial chips are selected to be common and evenly spaced, not to be functional. The overwhelming majority lie outside coding sequence, and there is no prior reason to expect the assayed base itself to alter a protein or its regulation.

Second, a marker is correlated with everything physically near it, because nearby loci are inherited together. An association therefore implicates a region, not a position. The size of that region is set by the extent of linkage disequilibrium in the population, which in livestock can be hundreds of kilobases and may contain many genes.

What the result supports is that a variant affecting milk yield lies somewhere in the neighbourhood of that SNP. Narrowing the neighbourhood to a gene requires denser markers, sequence data, or functional evidence, and is a substantial separate undertaking.

Exercise 1.4

Choosing a layer. You want to predict which young bulls will produce high-yielding daughters, and you may measure one 'ome'. Which do you choose, and what is the argument against choosing the transcriptome, which is closer to the phenotype?

Show solution ▾

Choose the genome. The argument is not that the genome is more informative about the phenotype in a biological sense, because the transcriptome is closer to it. The argument is about what can be measured, when, and in what tissue.

The genome is identical in every cell, does not change with age, physiology or season, and can be assayed from a blood or hair sample at birth. One measurement is valid for the animal's whole life. The transcriptome differs between tissues and changes continuously, so a useful measurement would require the right tissue at the right time, which for a lactating mammary gland is not available from a newborn bull.

Since the entire value of genomic selection lies in evaluating animals before they have records, a predictor that requires the animal to be adult and lactating defeats the purpose. This is the reasoning developed in Lecture 10.

Adapted by ASAP-Bio from the Introduction to Genomics slide series by Gabor Meszaros (Genomics Boot Camp), released under GPL-3.0. The sequence and worked examples follow his; the explanatory text was written for this platform.
Ministry of Foreign Affairs of Denmark Danida Fellowship Centre
The project is funded by the Ministry of Foreign Affairs of Denmark and managed by Danida Fellowship Centre.
DANIDA Knowledge and Innovation Programme (KIP) 2025.