Research

We study how genomic variation across Indian populations shapes health and disease, working across four connected themes: population genomics, epigenetics, microbial genetics, and algorithms and AI.

01

Population Genomics

India is home to some of the most linguistically, culturally, and genetically diverse populations on Earth, shaped by millennia of migration, isolation, and endogamous marriage practices within many communities. Yet global reference datasets and variant catalogues have historically underrepresented this diversity, built predominantly from populations of European ancestry. Our lab works to close that gap: building population-scale genomic resources for India, and using them to ask what patterns of variation reveal about demographic history, and about disease risk carried disproportionately by specific communities.

Structural Variants (SVs)

Structural variants (deletions, duplications, insertions, and larger genomic rearrangements) are a major source of functional genetic variation, but have been chronically underrepresented in population studies because they’re hard to detect accurately from short-read data alone. Using long-read sequencing, we’re building structural variant catalogues across Indian populations, filling a real gap in the global SV reference landscape. This has a direct connection to demographic history: several Indian communities show strong evidence of founder effects and sustained endogamy, and elevated homozygosity, including at structural variant loci, is one of the clearest genomic signatures of that history. Characterising SV burden at population scale lets us examine not just where these founder events happened, but what their consequences are for the deleterious variant load carried by descendant populations today: work that sits at the intersection of population history and clinical genetics.

Tandem Repeats (TRs)

Tandem repeats are disproportionately important to human disease relative to their share of the genome: dozens of neurodegenerative and neuromuscular disorders are caused by pathogenic repeat expansions, yet population-scale surveys of TR variation have lagged behind SNP-based studies considerably, largely because TRs are so difficult to genotype accurately. Applying our long-read TR genotyping methods (described under Algorithms and AI) across population cohorts, we’re characterising the normal range of TR variation in Indian populations and identifying pathogenic expansions directly in cohort data, work that matters both for understanding baseline population diversity at these loci, and for building a foundation for better-informed clinical interpretation of repeat expansions in Indian patients, who are currently assessed against reference ranges drawn almost entirely from other populations.

02

Epigenetics

Beyond the DNA sequence itself, we’re interested in how chemical modification of DNA and chromatin state regulate genome function, and in building better tools to actually observe these modifications directly, rather than inferring them indirectly.

A central focus is nanopore sequencing’s ability to detect base modifications directly from raw signal, without the bisulfite conversion or antibody-based enrichment that classical methylation assays require. We’ve built models for accurately identifying 6-methyladenine (6mA), a modification that, unlike the extensively studied 5-methylcytosine, has been poorly served by existing detection tools, despite its established role in bacterial epigenetic regulation and growing evidence of relevance elsewhere. Because this detection happens directly from sequencing signal, it opens up 6mA analysis in both new experiments and in existing nanopore datasets that were never generated with modification-calling in mind.

We’re also interested in where epigenetic regulation and tandem repeat biology intersect: methylation state at TR loci is itself a source of functional variation that’s rarely examined alongside repeat length or motif composition. Our visualisation tool VisuaMiTRa was built specifically to make allele-specific motif composition and methylation profile explorable side by side, rather than requiring separate analyses that are difficult to connect back to each other.

Separately, we’ve studied epigenetic reprogramming during early development, specifically how transient reactivation of a gene network normally restricted to the earliest embryonic (2-cell) stage facilitates mouse embryonic stem cells’ ability to self-organise into blastoid structures. This work speaks to a broader question in the lab: how transitions in chromatin state, not just static epigenetic marks, shape what a cell is capable of becoming.

03

Microbial Genetics

Our microbial genetics work centres on antimicrobial resistance (AMR), tracking it through genomic and metagenomic surveillance in two very different settings: individual clinical infections, and population-level environmental monitoring through wastewater.

On the clinical side, in collaboration with the LV Prasad Eye Institute, we conducted genomic surveillance of bacterial pathogens causing eye infections across India, combining whole-genome sequencing with antimicrobial susceptibility testing on 291 high-quality bacterial genomes. The study identified vancomycin-resistant Staphylococcus aureus and extensively drug-resistant (XDR) Klebsiella pneumoniae among the isolates, along with previously uncharacterised AMR-associated genetic mechanisms, findings with direct clinical relevance since ocular infections are typically treated empirically before laboratory susceptibility results are available, and the resistance patterns observed don’t always match what current treatment guidelines assume.

On the environmental side, we conducted a two-year metagenomic study of wastewater across 19 sites in four major Indian cities, treating municipal wastewater as a population-level surveillance tool for AMR that doesn’t depend on individual clinical sampling. The study found that microbial community composition varied substantially by city, but, more strikingly, the underlying resistome (the collection of antibiotic resistance genes present) was largely homogenous across cities regardless of that compositional difference, with genes conferring resistance to tetracyclines and beta-lactams showing a stronger association with mobile genetic elements than macrolide resistance genes, suggesting a differing capacity for horizontal spread. The analysis also recovered a substantial number of metagenome-assembled genomes with no close match in existing databases, pointing to how much of the microbial diversity in these environments remains uncharacterised.

Together, these two threads reflect a consistent approach: AMR isn’t just a clinical problem or an environmental one; the same resistance genes and mobile elements move between both settings, so surveillance that only looks at one side misses how resistance actually spreads through a population.

04

Algorithms and AI

We treat method and tool development as a research output in its own right, not merely infrastructure supporting the biology above; several of our most-used tools began from finding non-obvious computational structure in a problem that looked, on its face, like it demanded brute-force comparison.

Our tandem repeat work is a good example of this philosophy in practice. DiviSSR came from recognising that DNA sequences, represented in 2-bit form, can be tested for repetitiveness through simple arithmetic (whether a resulting number is divisible by, or leaves a specific remainder against, a particular value) rather than exhaustive pattern matching, letting it identify all repeats in a human genome in about 30 seconds on an ordinary laptop. PERF, an earlier tool in the same lineage, used hash-based membership testing to bring repeat identification in the human genome down to a couple of minutes, and remains widely used. Ribbit extends this numeric-representation approach to a harder problem: resolving complex, imperfect, and nested tandem repeat structures, handling indels and substitutions gracefully and decomposing compound repeat loci that simpler callers can’t correctly separate. And ATaRVa brings this lineage to long-read data specifically: a sequencing technology-agnostic genotyper that works across both PacBio and Oxford Nanopore platforms, runs roughly an order of magnitude faster than existing tools, and is specifically tuned to remain accurate in the low-coverage settings common in clinical sequencing, including correctly resolving pathogenic expansions that naive genotypers misclassify due to benign, interrupted alleles nearby.

Beyond tandem repeats, we build machine learning models for tasks where signal-level nanopore data is hard to interpret directly, including our 6mA detection models described under Epigenetics. We also treat rigorous benchmarking as its own research contribution: for example, examining whether modern long-read basecalling models introduce systematic artefacts specifically at repeat loci, a question with direct consequences for anyone using these tools in a clinical setting, where a confidently wrong genotype is worse than a flagged uncertain one. A fast, elegant tool that’s quietly inaccurate isn’t actually useful, so validating our own tools against ground truth, and against each other, is a standing part of how we work rather than a one-time step before publication.