A new handbook released by the Ministry of Statistics and Program Implementation (MoSPI) lays out a roadmap to standardise, connect and prepare government datasets for AI-powered governance.
The ministry says India already possesses one of the world’s richest administrative data ecosystems, including:
- 90 crore+ Ayushman Bharat health records
- 69 crore+ DigiLocker users
- 44 crore+ vehicle registrations
- 31 crore+ eShram workers
- 24.6 crore property records
- 9.2 crore land records
Yet these databases often use different formats, identifiers, definitions and classifications, making it difficult to combine information across departments.
As the handbook notes, abundant data does not automatically create intelligence. AI systems require structured, comparable and trustworthy information before they can generate meaningful insights.
How does the government plan to solve it?
The strategy is to harmonize data, not centralize it. Instead of creating a single, central database, ministries will continue to own their datasets while adopting common metadata, identifiers, classifications, quality standards and APIs.
This will enable government systems to exchange, interpret and reuse data consistently across departments.
Where does AI fit into the plan?
According to the handbook, AI comes only after data becomes reliable and interoperable.
“AI readiness begins with data readiness,” it states, warning that AI models trained on fragmented or poorly documented datasets could amplify inconsistencies instead of improving governance. Harmonized data, by contrast, can support trusted analytics and evidence-based policymaking.
Principal Secretary to the Prime Minister Pramod Kumar Mishra echoed this message at Statistics Day, saying the priority now is to standardize data across ministries, make it interoperable and derive trusted insights from it.
“A lot of data is generated by our departments, ministries and digital activities also. So, how to use the potential of the data, how to standardize it, how to make it compatible with each other, how to derive inferences from the data, how to ensure that the data is as trusted as comprehensive and as compatible as the earlier survey data or maybe it is expanding the scope of our statistics,” he told ANI.
How far has India come?
The handbook highlights India’s rapid expansion of digital public infrastructure through platforms such as Ayushman Bharat, DigiLocker, eShram, land records and vehicle registrations. While these systems have generated enormous volumes of data, the ministry believes the next stage is making these datasets interoperable so they can be securely reused across government.
What’s next?
MoSPI has proposed a three-phase roadmap. Departments will first document and organize existing datasets, then align them using common standards and quality checks, before making them discoverable through catalogs and APIs.
The long-term objective is to create datasets that are machine-readable, interoperable and reusable for AI applications.
What is the end goal?
The ministry envisions India’s digital journey evolving from Digital Public Infrastructure to a Harmonized Data Pipeline, and ultimately to a Public Intelligence Infrastructure, where trusted, AI-ready datasets support policymaking, public service delivery and population-scale analytics.
Overall: I’d rate this 9.5/10 after these edits. The only substantive change I’d recommend is replacing “fix its data” with “harmonise its data” in the headline where possible, as it better reflects the handbook’s language and avoids implying the existing data is flawed.




