Skip to content

Genetics: Population Diversity, Genome India Project & Rare Diseases

1. GENETIC INHERITANCE & POPULATION COMPLEXITY
Cue WordsNotes
DNA and the Architecture of Inherited Disorders
  • **DNA as the Genetic Blueprint**: Composed of four nucleotide bases — Adenine, Thymine, Guanine, Cytosine — whose codon arrangements transcribe the proteins executing cellular function.
  • **Single-Gene (Mendelian) Disorders**: Mutations at a single gene locus (e.g., Sickle Cell Anaemia, Cystic Fibrosis) — the most tractable targets for gene-editing therapies.
  • **Polygenic Disorders**: Arise from complex multi-gene and environmental interactions (e.g., hypertension, diabetes) — harder to correct via single-gene intervention.
  • **Core Glossary**: Base = compound bonding with another base to form DNA/RNA; DNA = double-stranded helical molecule carrying genetic instructions; DNA sequence = specific order of the four bases; Genome = an organism's complete set of DNA-encoded instructions; RNA = DNA's single-stranded cousin that produces proteins.
Consanguinity, Endogamy & Founder Effects
  • **Population Complexity**: India hosts **4,600+ distinct population groups** with high endogamy, restricting gene flow and concentrating recessive disease-causing genes within specific community clusters ("founder effects").
  • **Consanguinity Risk**: Marriages between close biological relatives raise the risk of autosomal recessive genetic disorders by **2x to 3x**.
  • **Rare Disease Burden**: India has an estimated **7.0 to 9.6 Crore rare-disease patients** (~6-8% of the population), of which **80%** trace to genetic origins — a direct consequence of this endogamy-driven mutation load.
> **Summary**: India's extreme endogamy across 4,600+ population groups creates localized founder mutations, translating high consanguinity rates directly into one of the world's largest rare-disease patient burdens.
2. THE GENOME INDIA PROJECT & POPULATION-SCALE MAPPING
Cue WordsNotes
Genome India Project: Scope & Infrastructure
  • **Completion Milestone**: The Department of Biotechnology (DBT) completed sequencing **10,000 reference genomes**, representing 99% of India's endogamous population groups.
  • **Indian Biological Data Centre (IBDC)**, Faridabad: Houses the resulting reference-genome database — India's first population-specific genomic repository.
  • **Sequencing Economics**: The cost of sequencing a full human genome has crashed from **$3 Billion** (Human Genome Project, 2003) to **under $200** in 2026, making population-scale genomics fiscally viable.
National & Clinical Significance
  • **Diagnostic Sovereignty**: Prevents clinical misdiagnoses arising from reliance on Western-skewed reference genomes, which poorly represent Indian genetic variation.
  • **Vulnerability Mapping**: Enables identification of genetic markers predisposing specific Indian sub-populations to conditions such as cardiovascular disease or atypical drug-metabolism traits — the evidentiary base for personalised medicine.
> **Summary**: The completed Genome India Project gives India a sovereign genomic reference grid, correcting for Western-population bias and laying the diagnostic foundation for population-tailored personalised medicine.
3. RARE DISEASES & PHARMACOGENOMICS
Cue WordsNotes
National Policy for Rare Diseases (NPRD) 2021
  • **Group 1**: Disorders amenable to one-time curative treatment (e.g., bone-marrow transplants) — eligible for up to **₹50 Lakh** direct financial support per patient (raised from ₹20 Lakh).
  • **Group 2**: Disorders requiring long-term, high-cost therapy (e.g., Gaucher's disease).
  • **Group 3**: High-cost disorders with no definitive cure, requiring registry-based and clinical-trial research support.
  • **2026 Expansion**: Centres of Excellence for rare-disease diagnosis and treatment have grown from 8 to 15 across states and UTs, alongside 50+ nationwide awareness workshops conducted through 2026.
Pharmacogenomics: Genotype-Guided Prescribing
  • **Core Idea**: Integrates an individual's genetic profile with drug prescriptions to select optimal medications and adjust dosage to metabolic rate, avoiding lethal Adverse Drug Reactions (ADRs).
  • **Link to Genome India**: Population reference data from the Genome India Project sharpens pharmacogenomic prediction for Indian-specific metabolic variants, rather than relying on extrapolated Western pharmacogenomic norms.
> **Summary**: NPRD 2021's tiered financial-support model, now backed by an expanding Centre-of-Excellence network, is converging with Genome India's population data to move India from reactive rare-disease relief toward pre-emptive pharmacogenomic prescribing.
UPSC Mains PYQs
  • Genomic Projects: Explain the objectives and national significance of the Genome India Project. Discuss how population-level genetic mapping can assist in developing personalised medicine and identifying genetic disease vulnerabilities in India. (10 Marks, 150 Words)
  • Rare Disease Policy: Critically examine India's National Policy for Rare Diseases, 2021, in light of its expanding Centre-of-Excellence network. Discuss the role of consanguinity and endogamy in shaping India's genetic disease burden. (10 Marks, 150 Words)
  • The 2024 Nobel Prize in Physiology or Medicine was awarded to Victor Ambros and Gary Ruvkun for discovering microRNA and its role in post-transcriptional gene regulation; their work in the nematode C. elegans (genome ~100 million base pairs) identified the lin-4 microRNA regulating the lin-14 mRNA by binding its untranslated region and blocking translation (not coding for protein itself) -- overturning the earlier view that transcription factors alone control gene regulation, and the related let-7 microRNA was later found conserved across animals including humans, with implications for evolution, disease/mutation understanding, stem cell biology, and cancer detection.
  • The genetic code consists of 64 codons (nucleotide triplets): 61 specify amino acids and 3 are stop codons; it is described as universal (same across organisms), degenerate (multiple codons can code for one amino acid), and non-overlapping/without "punctuation" -- Har Gobind Khorana's work was instrumental in deciphering this code, for which he shared the 1968 Nobel Prize in Physiology or Medicine.
  • A Single Nucleotide Polymorphism (SNP) is a naturally occurring variation in the DNA sequence where a single base (A, T, C, or G) differs between individuals at a specific genomic location; SNPs act as biomarkers and underlie much of human genetic diversity and disease-susceptibility research.
  • DNA mutations -- changes in an organism's DNA sequence -- are classified as deletion, duplication, inversion, or point mutation; sickle cell anemia is a classic example caused by a single point mutation in the beta-globin gene that alters hemoglobin structure, causing red blood cells to sickle.
  • Chromosomal disorders arise from abnormalities in chromosome number (aneuploidy, e.g., due to non-disjunction during meiosis) or structure. Numerical disorders include Down syndrome (Trisomy 21, an extra copy of chromosome 21), Turner syndrome (45,X -- a missing X chromosome, affecting females), and Klinefelter syndrome (47,XXY -- an extra X chromosome, affecting males). Structural disorders include deletion (a chromosome segment is lost, e.g., Cri-du-chat syndrome), duplication, translocation (a segment moves to a non-corresponding chromosome), and inversion (a segment breaks, flips, and reattaches).
  • India's Unified Genomic Chip is a single nucleotide polymorphism (SNP) chip enabling genomic profiling of dairy animals to support genetic improvement and productivity in India's livestock sector, with separate chip variants developed for cattle and buffaloes.
  • The Human Genome Project's goals were to: identify all genes in human DNA (an estimated 20,000-25,000 genes); determine the sequence of about 3.2 billion chemical base pairs; store this information in databases; improve tools for data analysis; transfer the resulting technology to industry; and address the ethical, legal and social issues (ELSI) arising from genome research.
  • Key findings of the Human Genome Project: the genome comprises about 3.2 billion base pairs distributed across 23 pairs of chromosomes; only 1-2% of DNA codes for proteins, while most non-coding DNA (once dismissed as 'junk') is now known to be functional, containing regulatory elements and repetitive sequences; humans have surprisingly few genes (~20,000), showing organismal complexity does not come from gene count alone but from 'alternative splicing' (a single gene read differently to produce multiple proteins); and humans are 99.9% genetically identical, with the remaining fraction (e.g. SNPs) accounting for human diversity.
  • The Telomere-to-Telomere (T2T) Project sequenced heterochromatin regions of the human genome that were not sequenced in the Human Genome Project and had been thought to be junk DNA; this is important for understanding genetic variation, how the genome works as a whole, genetic diseases, and pharmacogenetics.
  • Chromatin exists in two states: heterochromatin, which is more condensed, contains silenced (methylated) genes, is gene-poor with high AT content, and stains darker; and euchromatin, which is less condensed, contains actively expressing genes, is gene-rich with high GC content, and stains lighter -- a distinction relevant to gene regulation and epigenetics.
  • India's key genomic initiatives: the Indigen Project, conducted by CSIR, performs whole-genome sequencing of 1,000 individuals; the Genome India Project, led by the Department of Biotechnology (DBT), collects 10,000 genetic samples to build a reference genome and map India's genetic diversity, supporting the Indigen People's Initiative, which aims to improve identification and classification of variants linked to Mendelian disorders and support precision medicine.
  • Reverse transcriptase is an enzyme that synthesises Deoxyribonucleic Acid (DNA) from a Ribonucleic Acid (RNA) template, enabling the reverse transcription reaction central to retroviruses (e.g. HIV) and to lab techniques like RT-PCR.
  • Tmesipteris oblanceolata, a species of fork fern, holds the record for the largest known genome, containing about 160 billion base pairs -- outstripping the human genome (about 3.2 billion base pairs) by more than 50 times.
  • Cell-free DNA (cfDNA) refers to small fragments of nucleic acids that are released from cells and found outside the cell in body fluids such as plasma, urine, and cerebrospinal fluid (CSF), used in applications like non-invasive prenatal testing and liquid-biopsy cancer detection.
  • The mitogenome is a small, circular chromosome present inside mitochondria, made of double-stranded DNA like the nuclear genome, but unlike nuclear DNA it is inherited only from the mother (maternal/matrilineal inheritance).
  • Selective silencing is a process in which a cell or organism specifically turns off the expression of certain genes, either naturally or through intervention, to control cellular functions, study gene roles, or treat diseases -- often by inactivating just one parent's gene copy (genomic imprinting) or by using RNA interference (RNAi) to block harmful messenger RNA.
  • DNA methylation involves the addition of a methyl group to the DNA molecule; high levels of DNA methylation lead to gene silencing, making it a key mechanism studied in epigenetics (how behaviour and environment can cause changes affecting gene function without altering the DNA sequence).
  • Cis-regulatory elements (CREs) are non-coding DNA sequences, located near the genes they control, that act as switches to precisely regulate when, where, and how much a gene is turned on (transcription), functioning as crucial parts of genetic networks guiding development, adaptation, and disease.
  • Next-Generation Sequencing (NGS) is a modern method of analysing genetic material that can rapidly sequence large amounts of DNA or RNA; NGS can sequence an entire genome within days, compared to months with earlier sequencing techniques.
  • Extrachromosomal DNA (ecDNA) refers to small circular DNA fragments that are present inside the nucleus but separate from the chromosomes; ecDNA has been linked to how cancer cells amplify oncogenes and drive tumour evolution.