Methods & data
How the numbers on this site are made โ and their limits. Last reviewed 2026-07-12.
How we find the blind spots
D2I2 catalogues 1,648 diseases that matter in India and asks a simple question of each one: is this a place where a genomics study could actually move the needle for Indian patients, and is anyone looking?
We answer in two tiers. Tier 1 is the tractable frontier: 22 diseases where there is a real genomic handle in our data. A handle means one of three concrete things - a European-built risk score whose mis-reading of South Asians we can measure, a gene set with pathogenic variants actually seen in South Asians, or a single variant that is far more common here than in Europe. These get a number. Tier 2 is 276 high-burden diseases with no verified India incidence and no genomic handle yet. We do not dress these up as ready studies; we mark them honestly as data gaps.
The Tier 1 number weights three things: how much the disease burdens India, how big the South-Asian blind spot is, and how tractable a study would be. The gap and the burden carry most of the weight because a blind spot only matters where the disease is common.
Technical detail
Tier-1 whitespace score = 0.45 x IndiaBurden + 0.45 x SouthAsianGap + 0.10 x Tractability, each term on a 0-1 scale.
IndiaBurden reflects India disease burden; SouthAsianGap is the size of the measured South-Asian genomic blind spot (PRS mis-stratification magnitude, or the count of South-Asian-observed pathogenic variants relative to the strongest gene set); Tractability reflects how runnable a validation study is once an Indian genotyped-plus-phenotyped sample is in hand. Burden and gap are weighted equally and dominate, because a transferability gap is only worth chasing where the disease is common. Tractability is a light thumb on the scale, not a driver. A disease qualifies for Tier 1 only if it carries at least one of the three handle types (prs, common_variant, rare_variant); everything else that is high-burden but handle-less falls to Tier 2 as a descriptive gap. The score deliberately concentrates on the tractable frontier because you can only run a genomics study where a handle exists.
How a risk score can misread an ancestry
A polygenic risk score adds up many small-effect DNA variants into one number that estimates your genetic risk for a disease. The catch: almost every score in wide use was built and calibrated on European data. Its cut-off for 'high risk' is set to the European top 10 percent.
Think of it as a ruler set for the wrong crowd. If you calibrate a height ruler on one population and then hold it up against another whose average is different, everyone in the second group reads too tall or too short - not because they changed, but because the ruler's zero is in the wrong place. That is what a European risk score does to South Asians.
We measure the error two ways. 'Mis-stratification percent' is the share of a population that lands above the European high-risk cut-off. Europeans sit at 10 percent by definition. When a group reads 20 percent, the score is flagging twice as many people as it was designed to - not because they are sicker, but because the ruler is off. 'Mean-shift in SD' says how far that group's average score sits from the European average, measured in standard deviations. A shift of +0.50 SD means the whole distribution has slid up by half a standard deviation.
Worked example - the Type 2 diabetes over-flag. We ran the European-built T2D score PGS000033 two ways. On GenomeIndia's pooled ~10,000-genome all-India sample it places 20.7 percent of Indians above the European high-risk cut-off - roughly twice the 10 percent it was designed for, a +0.50 SD upward mean-shift. The smaller, diaspora-sampled 1000 Genomes Telugu group reads sharper still at 30.6 percent (+0.80 SD). Both say the same thing at different resolutions. One reason is the MTNR1B risk variant rs10830963, carried at about 40 percent in India (GenomeIndia) and 43 percent in the 1000 Genomes South-Asian panel, against 29 percent in Europeans. The score is not wrong about the DNA; its zero point is simply set for the wrong crowd. Note this cuts both ways: for coronary artery disease the same method shows the European score placing only 2 percent of the pooled all-India sample above the cut-off, badly UNDER-warning a population with well-documented excess heart disease.
Technical detail
Method: analytical Hardy-Weinberg mean/variance from allele frequencies. For a PRS summed over independent variants, the population mean is sum(2 x p_i x beta_i) and the variance is sum(2 x p_i x (1 - p_i) x beta_i^2), where p_i is the effect-allele frequency in that population and beta_i the reported effect size. Swapping the EUR frequency vector for a target population's vector shifts the mean; we express that shift in EUR standard-deviation units (shift_vs_eur_sd). Mis-stratification percent is then the mass of a Normal(shifted mean, population variance) lying above the EUR 90th-percentile threshold; EUR is 10.0 percent by construction. Two frequency substrates are used side by side. The per-subpopulation rows use 1000 Genomes phase-3: for PGS000033 the ordered shifts run EUR 0.0 SD (10.0%), through PJL +0.544 SD (22.4%), GIH +0.587 (22.6%), BEB +0.641 (24.5%), STU +0.789 (29.9%), to ITU +0.796 SD (30.6%). The pooled All-India row uses GenomeIndia (~10,000 Indian genomes, GRCh38): for PGS000033 that is +0.50 SD / 20.7% (computed over the scoring variants GenomeIndia covers, against a EUR baseline restricted to the same set). GenomeIndia's HWE/inbreeding QC drops a few high-Fst common variants that fail Hardy-Weinberg in a pooled multi-population sample (a Wahlund effect), so some scoring variants have no All-India value; the count used per score is recorded in the data. This model captures the allele-frequency mean-shift ONLY. It does not model linkage-disequilibrium differences or effect-size (beta) heterogeneity between ancestries, both of which can further degrade transferability. It is an analytical screen for where mis-calibration is likely and large, not an individual-level validated risk estimate.
Our data sources & versions
Every number on D2I2 traces to a public, citable source. We do not generate genotypes; we combine reference datasets to spotlight where an India-facing study is missing.
The pooled all-India allele frequencies - the 'All-India' bar and PRS row - come from GenomeIndia, the national project that released whole-genome data for about 10,000 Indians (GRCh38) in January 2025. That is the large-sample number. Alongside it we keep the 1000 Genomes Project phase 3 view, which is smaller but breaks down into five South-Asian sub-populations: Punjabi from Lahore (PJL), Gujarati (GIH), Bengali (BEB), Indian Telugu (ITU) and Sri Lankan Tamil (STU), each around 86 to 103 samples, roughly 489 South Asians in total. We keep 1000 Genomes for that finer per-community resolution, which GenomeIndia's open release does not provide - it publishes a single pooled all-India frequency, not per-community numbers. Polygenic scores are pulled by their stable PGS Catalog IDs (for example PGS000033 for type 2 diabetes). Variant pathogenicity predictions come from AlphaMissense, accessed through dbNSFP and MyVariant.info. Where-is-it-seen and how-often counts come from gnomAD. Plain-language disease summaries are drawn from Wikipedia and then simplified.
Technical detail
GenomeIndia (IBDC/RCB, released Jan 2025; ~10,000 Indian whole genomes, GRCh38; ~129.9M variants after QC) - the pooled all-India allele-frequency summary, joined on chr:pos:ref:alt (the open file carries no rsIDs) and mapped onto each panel/scoring effect allele with strand handling. It supplies the 'All-India' (IN) column in the frequency table and the pooled All-India row in the PRS mis-stratification. GenomeIndia's open release is a SINGLE pooled all-India frequency; there is no per-subpopulation (Telugu/Punjabi/etc) breakdown on the open tier, so we do not claim one. We publish only the derived per-variant numbers we extract, with a GenomeIndia citation, and do not re-host the raw file. 1000 Genomes Project phase 3 (2015 release, GRCh38-mapped allele frequencies) - South-Asian panel PJL (n~96), GIH (n~103), ITU (n~102), STU (n~102), BEB (n~86); ~489 SAS individuals; note GIH/ITU/STU are diaspora-sampled (Houston/UK); retained for per-subpopulation resolution the GenomeIndia open file lacks. PGS Catalog - scores referenced by immutable PGS IDs (PGS000033, PGS000010, PGS000047, PGS000061, PGS000067, PGS000350, etc.), harmonised GRCh38 positions used to look up GenomeIndia frequencies. AlphaMissense (Cheng et al., Science 2023; ~71M proteome-wide missense predictions) served via dbNSFP and the MyVariant.info API. gnomAD (v4-era joint exome+genome frequencies) for South-Asian observed counts and European-absence flags. Wikipedia for lay summaries, cleaned of stray replacement characters before display. All version pins are advisory: the pipeline that regenerates these JSON files is the source of truth for the exact build vintage.
What we do NOT claim (limitations)
D2I2 is a map of where to look, not a diagnosis and not a finished result. Three honest limits.
First, the PRS mis-stratification is an analytical estimate from allele frequencies. It has NOT been validated against real health outcomes in Indian individuals. It tells you a European score is probably mis-calibrated here and roughly how badly; it does not tell you any one person's risk.
Second, the reference panels are small. A few hundred South Asians in 1000 Genomes, mostly sampled from the diaspora, cannot capture the genetic depth of a subcontinent with thousands of endogamous communities. A signal absent from the panel is not proven absent from India.
Third, our model captures the allele-frequency shift and nothing more. Real transferability also breaks on linkage-disequilibrium differences and on effect sizes that differ between ancestries. Those can make the true gap larger or smaller than our screen suggests. Every Tier 1 entry is a fundable hypothesis, not a conclusion.
Technical detail
The mean-shift model assumes Hardy-Weinberg equilibrium, variant independence (no LD modelling), and ancestry-invariant effect sizes - all approximations. Consequences: (1) shift_vs_eur_sd and misstrat_pct are screening estimates, unvalidated against individual Indian phenotype data; (2) small, partly-diaspora reference panels under-sample India's endogamous structure, so European-absence and SAS-observation flags carry sampling error; (3) LD and beta heterogeneity are unmodelled, so true PRS transferability loss may exceed or fall short of the allele-frequency component alone. Tier-1 items are validation hypotheses conditional on obtaining a genotyped-plus-phenotyped Indian cohort; they are not claims about deployed clinical performance.
Data provenance: GenomeIndia is now the pooled all-India number
For a long time our allele frequencies rested only on a few hundred South Asians in the 1000 Genomes reference panel - thin ground for a country of India's genetic depth. The pooled part of that gap is now closed.
In January 2025 the GenomeIndia project released whole-genome data for about 10,000 Indians spanning dozens of communities (GRCh38), archived at the Indian Biological Data Centre (IBDC) as a public, no-login allele-frequency summary. D2I2 now uses it directly: the 'All-India' bar on each disease and the pooled All-India row in the PRS tables are computed on GenomeIndia's ~10,000-genome frequencies, not on 1000 Genomes. That answers the obvious objection - 'you trust a few dozen samples?' - for the pooled number. The type 2 diabetes over-flag and the coronary-artery under-flag both hold on the India-native data.
One honest limit remains, by design. GenomeIndia's open release is a single pooled all-India frequency - it does not break down into Telugu, Punjabi, Gujarati, Bengali or Tamil. That per-community resolution sits behind an institutional managed-access gate (IBDC's FeED protocol, which needs a Govt/academic affiliation) that an independent project cannot pass. So we keep the 1000 Genomes per-subpopulation view alongside GenomeIndia, clearly labelled as the finer-resolution but much smaller-sample source. Getting per-community India-native frequencies is a job for an academic partner (for example CCMB or Plaksha) with managed access - it is on the roadmap, not something we fake from the open file. Two related resources round out the picture: CSIR's IndiGen programme (1,029 Indian genomes, IndiGenomes/CSIR-IGIB) and GenomeAsia 100K (a 1,739-individual pan-Asian pilot, Nature 2019).
Technical detail
Pooled all-India substrate (live): GenomeIndia open allele-frequency summary - ~10,000 Indian whole genomes, GRCh38, ~129.9M variants post-QC, IBDC/RCB, released Jan 2025 under FeED / BIOTECH-PRIDE Open-Access. Joined on chr:pos:ref:alt (no rsIDs in the file), effect allele mapped with ref/alt + strand handling; PRS scoring positions taken from the PGS harmonised GRCh38 files. It supplies the IN column and the pooled All-India PRS row; per-score coverage is limited by GenomeIndia's HWE/inbreeding QC (a few high-Fst common variants fail HWE in the pooled multi-population sample - a Wahlund effect - and are dropped), and the count actually used is recorded per score. Per-SUBPOPULATION substrate (unchanged): 1000 Genomes phase-3 SAS (~489 individuals; PJL/GIH/BEB/ITU/STU), retained because the GenomeIndia OPEN tier has no per-community breakdown - that lives behind IBDC managed access (institutional/Govt email required), which is an academic-partner deliverable (e.g. Plaksha/CCMB), not attempted here. Related cross-checks: IndiGenomes (1,029 genomes, CSIR-IGIB; ~32% India-unique variants) and GenomeAsia 100K (1,739-individual pilot, Nature 2019). Still outstanding: a versioned, DOI-tagged release pinning each build; and, critically, all of these remain analytical allele-frequency mean-shifts, not validated against Indian disease outcomes (that needs a genotyped-plus-phenotyped Indian cohort, e.g. UK Biobank's British South Asians or an India-native equivalent).
Sources
- GenomeIndia โ pooled all-India allele frequencies (IBDC download portal) โ
- GenomeIndia โ Mapping genetic diversity with the GenomeIndia project (Nature Genetics 57:767-773, 2025) โ
- Indian Biological Data Centre (IBDC) โ
- 1000 Genomes Project (phase 3) โ per-subpopulation South-Asian panel โ
- PGS Catalog (Polygenic Score Catalog) โ
- gnomAD (Genome Aggregation Database) โ
- AlphaMissense (Cheng et al., Science 2023) โ
- dbNSFP โ
- MyVariant.info โ
- IndiGenomes (CSIR-IGIB) โ
- GenomeAsia 100K โ
How to cite
D2I2: Decoding Disease in India. India Whitespace Score and genomic transferability atlas. d2i2.org (accessed 2026-07-12). Pooled all-India allele frequencies derived from GenomeIndia (Nat Genet 2025).
The pooled All-India frequencies are derived from the GenomeIndia open summary release (please also cite GenomeIndia / Nat Genet 2025 if you reuse those numbers); per-subpopulation frequencies are from 1000 Genomes phase 3. A versioned, DOI-tagged D2I2 release is planned so analyses can pin to an exact build vintage rather than a live-site snapshot.