D2I2.
Start here

New to all this? You’re exactly who we built it for.

You don’t need to know any biology to understand this site. Every hard word is underlined — tap it and a plain explanation pops up. Turn on student mode and the pop-ups lead with the simplest version.

A curious student looking up at a glowing DNA helix and the outline of a human body.

1. Start with the body

Pick a part of the body. You’ll see what it does and what can go wrong with it — heart, lungs, brain, gut, blood, and more.

2. Diseases you’ve probably heard of

Familiar names are a good place to begin — you already have a feel for them.

3. The big idea behind D2I2

Every cell in your body carries a set of instructions written in DNA. A gene is one instruction. Small differences in these instructions, called a , are part of why people are different from each other.

Here is the problem D2I2 is about. Almost all the science that links genes to disease was done on people whose families come from Europe. Which gene versions are common depends on , so a risk score or a test built for one group can quietly misread another.

South Asians are nearly a quarter of the world's people but a tiny slice of this research. That gap is the blind spot. Mapping it, and helping close it, is the whole point of this project.

4. Want to be a scientist?

These questions genuinely don’t have answers yet.

Real research isn’t only in textbooks. Here are open problems where a South-Asian study would matter — the kind of thing you could work on one day.

Type 2 diabetes
A European-trained polygenic risk score for type 2 diabetes flags 30.6% of Telugu (India) people as high-risk - vs the 10% it was designed for. That's a 3.1x mis-stratification: the score's 'average' is set to European genetics, so it systematically mis-reads South Asians (a +0.80 SD mean shift).
Cardiomyopathy
Across the cardiomyopathy gene set (LMNA, MYBPC3, MYH7, TNNI3, TNNT2), 150 AlphaMissense-'pathogenic' missense variants are actually seen in South Asians (gnomAD) - many European-absent and still clinically 'uncertain'. For cardiomyopathy, that's a pool of computationally-damaging, India-relevant, clinically-unresolved variants no one has systematically characterised.
Wilson's disease
Across the other high-burden (india) gene set (ATP7B), 75 AlphaMissense-'pathogenic' missense variants are actually seen in South Asians (gnomAD) - many European-absent and still clinically 'uncertain'. For wilson's disease, that's a pool of computationally-damaging, India-relevant, clinically-unresolved variants no one has systematically characterised.
Coronary artery disease
A European-trained polygenic risk score for coronary artery disease places only 0.7% of Sri Lankan Tamil people above its high-risk cut-off - far BELOW the 10% it was calibrated to. Yet South Asians carry a well-documented EXCESS of real-world heart disease. The score is blind to South-Asian coronary artery disease genetics: it under-warns exactly the group at higher true risk. This under-flagging is more dangerous than over-flagging.
Dilated cardiomyopathy
Across the cardiomyopathy gene set (LMNA, MYBPC3, MYH7, TNNI3, TNNT2), 150 AlphaMissense-'pathogenic' missense variants are actually seen in South Asians (gnomAD) - many European-absent and still clinically 'uncertain'. For dilated cardiomyopathy, that's a pool of computationally-damaging, India-relevant, clinically-unresolved variants no one has systematically characterised.
Hypertrophic cardiomyopathy
Across the cardiomyopathy gene set (LMNA, MYBPC3, MYH7, TNNI3, TNNT2), 150 AlphaMissense-'pathogenic' missense variants are actually seen in South Asians (gnomAD) - many European-absent and still clinically 'uncertain'. For hypertrophic cardiomyopathy, that's a pool of computationally-damaging, India-relevant, clinically-unresolved variants no one has systematically characterised.

Be the scientist

Every red flag on this site is an open question waiting for someone to answer it. That someone could be you.

A genomics researcher spends real days doing this: pulling DNA data from public databases, writing a bit of code to compare how a gene variant shows up in one population versus another, spotting a pattern nobody has explained, then designing an experiment or a study to test it. Part detective, part coder, part biologist. You are looking for the thing that does not fit, and asking why.

If this pulls at you, take Biology seriously, and pick up some coding and statistics along the way. A genome is a giant dataset; the people who can both read biology and wrangle data are the ones who make discoveries. You do not need to have it all figured out at 15. Curiosity and the willingness to keep asking are the whole job.

And here is the honest pitch: the questions on D2I2 have no answers yet. The Indian data barely existed until recently. That is not a closed field you are joining late. It is a wide-open one, and you could be the person who solves a piece of it.

Indian Institute of Science (IISc) · BengaluruNational Institute of Biomedical Genomics (NIBMG) · Kalyani, West BengalCSIR-Institute of Genomics and Integrative Biology (CSIR-IGIB) · New DelhiCSIR-Centre for Cellular and Molecular Biology (CCMB) · HyderabadNational Centre for Biological Sciences (NCBS) · BengaluruInstitute for Stem Cell Science and Regenerative Medicine (inStem) · Bengaluru

Start with Biology plus a little code and statistics. Pick one disease on this site, ask why the Indian answer is missing, and follow it. The field is young enough that a student today can help write it.