How district data is reconciled
Census 2011 and NFHS-5 don't use the same district names, and both predate several districts that exist today. This page documents how each source is matched onto the app's current 785-district geography — including where that method is weakest.
The target geography
Every layer resolves to the same 785 districts, sourced from the Local Government Directory boundary releaserather than any one source's own snapshot. That release is itself dated — a small number of districts created since (in Madhya Pradesh, Arunachal Pradesh, Andhra Pradesh, Goa and Ladakh) aren't yet in it, and are mapped to their parent district in the meantime.
Matching by name, not by code
Census and NFHS-5 district names are matched to the current district list directly where the spelling already agrees, which clears the large majority of rows on its own. The rest go through a hand-built alias table — renamed districts (Gurgaon → Gurugram), transliteration differences, and reordered names — rather than a fuzzy best-guess match. A name that still doesn't resolve is left unmatched rather than guessed at.
Districts created after their source was published
Where a district didn't exist when Census 2011 was taken, its historical predecessor is determined geometrically — by intersecting the current district's boundary against 2011's — rather than by hand. Counts (population, households) are then apportioned across the successor districts by land area; rates and survey estimates are inherited unchanged from that predecessor instead, since a rate can't be meaningfully split by area. Every value produced this way is flagged in the data, not shown as an original measurement.
What's actually loaded
Census 2011 contributes population, density, literacy and 15 more demographic fields. NFHS-5 — a survey with roughly 115 published indicators — contributes the 7 with the most complete district coverage; the rest weren't imported rather than backfilled with a weaker estimate. Both are one-time imports, not a live feed, so a change upstream won't appear here until the app is reseeded.
Limitations
- Apportioning a parent district's counts by land area assumes population is spread evenly across it, which it rarely is — this affects only the districts created after 2011, not the layer as a whole.
- A district that inherits its predecessor's rate is identical to it on that indicator, which understates real variation between the two if treated as independent observations.
- NFHS-5 is a sample survey, not a census — district-level figures carry sampling error that isn't currently shown, and are least precise for the smallest districts.
- The highway network's total length figure double-counts stretches built as dual carriageways, since each carriageway is a separate segment in the source data.