Abstract
This tutorial is a learning resource that outlines the basic process and provides specific software tools for implementing a complete genome-wide association analysis. Approaches to post-analytic visualization and interrogation of potentially novel findings are also presented. Applications are illustrated using the free and open-source R statistical computing and graphics software environment, Bioconductor software for bioinformatics and the UCSC Genome Browser. Complete genome-wide association data on 1401 individuals across 861,473 typed single nucleotide polymorphisms from the PennCATH study of coronary artery disease are used for illustration. All data and code, as well as additional instructional resources, are publicly available through the Open Resources in Statistical Genomics project: http://www.stat-gen.org.
| Original language | English (US) |
|---|---|
| Pages (from-to) | 3769-3792 |
| Number of pages | 24 |
| Journal | Statistics in Medicine |
| Volume | 34 |
| Issue number | 28 |
| DOIs | |
| State | Published - Dec 10 2015 |
| Externally published | Yes |
Keywords
- Ancestry
- Bioconductor
- Call rate
- Genome-wide association (GWA) study
- Hardy-Weinberg equilibrium (HWE)
- Heatmap
- Heterozygosity
- IBD
- Imputation
- Lambda statistic
- Manhattan plot
- Minor allele frequency (MAF)
- Parallel processing
- Principal component analysis (PCA)
- Q-Q plot
- R code
- Regional association plot
- Relatedness
- Sample filtering
- SNP filtering
- Statistical genomics
- Substructure
- Tutorial
- UCSC Genome Browser
ASJC Scopus subject areas
- Epidemiology
- Statistics and Probability
Fingerprint
Dive into the research topics of 'A guide to genome-wide association analysis and post-analytic interrogation'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS