Ancient DNA + Random
Natural Selection
Disease Mapping
Genetic Variation
100

What is unique about ancient DNA? How can you distinguish between ancient DNA and present-day DNA?

Shorter reads (~50 bp avg length)

“Smiley-face” effect: high density of A/T at ends of reads due to C → U deamination

100

What does it mean if s = 1 for a disease-causing allele? What if s < 1? What if s = 0?

S = 1 indicates strong selection against the disease causing allele. If S < 1 there is still selection acting against the allele, but less so. If S = 0, there is no selection against the allele

100

What does an odds ratio = 1 indicate in a genetic association study?

That SNP has no association with that trait.

100

Which population of humans has the highest genetic diversity? Why?

African populations have the highest genetic diversity- this is because the remaining populations of humans originated from a bottleneck of early modern humans out of Africa (Out of Africa Migration).

200

What can we learn from ancient DNA that we cannot learn from present-day genomes alone?

A time-stamped catalog of variation, for reconstructing population history and tracing allele frequencies through time; Evaluate evidence for natural selection; Reconstruct genomes of ancestors who are extinct and understand how they relate to us; Investigate changes in evolutionary processes like mutation rate and recombination rate over time.

200

Why are some disease-causing alleles more common than you would expect?

Heterozygote advantage, bottleneck effect, founder events, small population size

200

What is SNP imputation and why is it important? What is the relationship between linkage disequilibrium and SNP imputation?

SNP imputation is when you leverage LD to impute untyped SNPs (basically find the identity of SNP based on what SNPs you already know that are nearby). This only works if the untyped SNP and the known SNP are in LD with each other (meaning they are linked).

200

Why do we use a factor of 2 when calculating divergence time? d=2𝜇T

To denote that there are 2 lineages that mutations are occurring on.

300

Where do the p-values in a GWAS Manhattan plot come from? 

Individual chi-square association tests for each SNP with the trait of interest.

300

Say you have identified a region of extended homozygosity in a population, indicative of positive selection. How would you figure out which locus in this haplotype is being selected for?

Look at nucleotide diversity across the region of homozygosity, do functional annotation of the region (are there relevant genes? What are regulatory annotations? Any variants of interest?), look at GWAS catalogs, functional validation (cell lines, model organisms)

300

How is a polygenic risk score calculated and what does it tell us?

A polygenic risk score is a sum of an individual’s genome-wide genotypes (SNPs) weighted by their effect size, as determined by GWAS data. Broadly polygenic risk scores inform us about the risk of having the trait and/or the severity of the trait given your genome- this has been applied to disease risk prediction.

300

What is the PRDM9 gene? Why is it relevant?

PRDM9 positions recombination hotspots across most mammals (>98% of hotspots). It works by binding DNA at specific locations to deposit histone methyl marks, which recruit recombination machinery.

400

Neanderthal gene flow into early modern humans happened ~50,000 years ago. Who has longer segments of Neanderthal ancestry in their genome, an individual who lived in Europe 30,000 years ago or an individual living in Europe today? Explain your answer.

An individual living 30,000 years ago would have longer segments of Neanderthal ancestry in their genome compared to a modern human because recombination has not had time to break up the Neanderthal haplotypes in this individual.

400

How would you distinguish a hard selective sweep from a soft selective sweep? What about adaptive introgression?

A hard sweep is where a beneficial allele, and the haplotype it is on, is quickly fixed in the population. A soft sweep shows multiple haplotypes rising in frequency. Adaptive introgression is when you can trace the swept haplotype back to interbreeding with another population or species (ex: Neanderthals).

400

After identifying a SNP associated with the trait of interest, what would you do next?

Fine-mapping (testing more SNPs in a smaller region), look for nearby genes (not that this will not always yield something), functional testing (cell lines, animal models), compare to regulatory and transcriptomic maps of the genome.

400

Why do we use a factor of 4 when calculating diversity (𝞹 = 𝞱 = 4Ne𝝁)?

2*Ne = we are diploid organisms, so a population of Ne has 2*Ne gene copies.

2*𝝁 = mutations are occurring on two lineages that we must account for.

500

What kinds of variants are best informed by linkage mapping in a pedigree? What about a GWAS?

A pedigree analysis would perform the best for simple traits with rare, large-effect variants (ex: a Mendelian trait). A GWAS best identifies common variants with small effect sizes for the trait of interest. This works best for complex or polygenic traits.

500

You sample a population of mice and find that it is not in Hardy Weinberg Equilibrium for a locus controlling coat color. There appears to be some positive selection for the gray color allele. What will the genotype frequencies be after 1 generation of random mating? Explain your answer.

The genotype frequencies would be Hardy Weinberg expected genotype frequencies. All it takes is 1 generation of random mating, regardless of parental genotypes, to reach HWE. Allele frequencies will remain the same, so you just need to calculate expected genotype counts.

500

What is population structure? Why would it confound a GWAS study?

Population structure is the idea that the population being studied is actually a mix of different sub-populations (ex: different ancestries). This can confound GWAS because if allele frequencies differ between sub-populations (due to drift, ancestry, etc.) and the trait of interest also differs between sub-populations, you will get spurious variants associating with the trait when they are actually just indicative of ancestry.

500

What is Lewontin's Paradox? How can we explain it?

Lewontin's Paradox is the observation that genetic diversity should increase with greater population size, but we observe in the field that this is not the case. This paradox can be explained by bottlenecks in populations, decreased diversity due to selection, and propagule size (more parental investment = less diversity).