Population Genomics
India is home to some of the most linguistically, culturally, and genetically diverse populations on Earth, shaped by millennia of migration, isolation, and endogamous marriage practices within many communities. Yet global reference datasets and variant catalogues have historically underrepresented this diversity, built predominantly from populations of European ancestry. Our lab works to close that gap: building population-scale genomic resources for India, and using them to ask what patterns of variation reveal about demographic history, and about disease risk carried disproportionately by specific communities.
Structural Variants (SVs)
Structural variants (deletions, duplications, insertions, and larger genomic rearrangements) are a major source of functional genetic variation, but have been chronically underrepresented in population studies because they’re hard to detect accurately from short-read data alone. Using long-read sequencing, we’re building structural variant catalogues across Indian populations, filling a real gap in the global SV reference landscape. This has a direct connection to demographic history: several Indian communities show strong evidence of founder effects and sustained endogamy, and elevated homozygosity, including at structural variant loci, is one of the clearest genomic signatures of that history. Characterising SV burden at population scale lets us examine not just where these founder events happened, but what their consequences are for the deleterious variant load carried by descendant populations today: work that sits at the intersection of population history and clinical genetics.
Tandem Repeats (TRs)
Tandem repeats are disproportionately important to human disease relative to their share of the genome: dozens of neurodegenerative and neuromuscular disorders are caused by pathogenic repeat expansions, yet population-scale surveys of TR variation have lagged behind SNP-based studies considerably, largely because TRs are so difficult to genotype accurately. Applying our long-read TR genotyping methods (described under Algorithms and AI) across population cohorts, we’re characterising the normal range of TR variation in Indian populations and identifying pathogenic expansions directly in cohort data, work that matters both for understanding baseline population diversity at these loci, and for building a foundation for better-informed clinical interpretation of repeat expansions in Indian patients, who are currently assessed against reference ranges drawn almost entirely from other populations.