Genomics
The International HapMap Project was a collaborative effort (2002–2009) that catalogued genetic variation and haplotype patterns across human populations, providing a public resource that revolutionized genome-wide association studies (GWAS) and modern medical genetics.1 Its name derives from "haplotype map," reflecting its goal of mapping blocks of linked variants inherited together.
The HapMap Project was launched in October 2002 as a natural successor to the Human Genome Project, aiming to provide a genome-wide map of common single-nucleotide polymorphisms (SNPs) and their haplotype structure.1 The rationale was that adjacent SNPs are often inherited together as haplotypes, so a subset of "tag SNPs" could capture most common variation, reducing the cost of association studies.2
The project was coordinated by an international consortium including the U.S. National Institutes of Health, the Wellcome Trust, and agencies in Japan, China, Canada, and Nigeria. Phase I (2005) genotyped 1.1 million SNPs in 270 individuals from four populations: Yoruba in Ibadan, Nigeria (YRI); Utah residents of Northern and Western European ancestry (CEU); Han Chinese in Beijing (CHB); and Japanese in Tokyo (JPT).1 Phase II (2007) added over 2 million more SNPs, and Phase III (2009) expanded to 11 populations and 1,184 individuals, including African, Asian, and admixed American groups.3
The HapMap resource enabled the design of genome-wide association studies by providing a reference panel for imputing untyped SNPs and for selecting tag SNPs.2 It directly facilitated the discovery of hundreds of genetic variants associated with common diseases, such as type 2 diabetes, coronary artery disease, and inflammatory bowel disease.4
Beyond disease mapping, HapMap data informed studies of human demographic history, natural selection, and recombination hotspots. For example, the project's dense SNP data revealed that recombination is concentrated in narrow hotspots, and that patterns of linkage disequilibrium vary across populations, reflecting differences in effective population size and migration.5 The project also set ethical standards for community engagement, including the use of population-specific consent and the establishment of a data access policy that balanced openness with privacy.3
Although the HapMap Project officially ended in 2009, its data remain a foundational resource, and its approach was extended by the 1000 Genomes Project, which provided a more comprehensive catalog of variants, including rare and structural variants.6 The HapMap's tag SNP concept was instrumental in the design of commercial genotyping arrays, such as those from Illumina and Affymetrix, which were widely used in GWAS.4
HapMap also influenced the development of the International Genome Sample Resource (IGSR), which maintains and distributes the original HapMap cell lines and data, ensuring continued access for researchers.3 The project's emphasis on public data release and community norms set a precedent for later large-scale genomics initiatives, including the Human Cell Atlas and the All of Us Research Program.
One overlooked aspect is the project's role in developing statistical methods for haplotype phasing and imputation, such as the PHASE algorithm and the IMPUTE software, which were later refined for whole-genome sequencing data.5 Another is the inclusion of the Luhya in Webuye, Kenya, and the Maasai in Kinyawa, Kenya, in Phase III, which were chosen to represent East African diversity, but the project was criticized for not including more African populations given the continent's high genetic diversity.3
Ethically, the project faced debates over the commercialization of cell lines derived from donor samples, leading to a revised consent that allowed for limited commercial use, but with restrictions on patenting.3 The HapMap also pioneered the use of "community engagement" models, where researchers consulted with donor communities about data use, a practice that became standard in later projects like the 1000 Genomes Project.6
The HapMap Project's data are available through the International Genome Sample Resource (IGSR) at https://www.internationalgenome.org.
Help improve the encyclopedia. Reports go straight to the site manager.