The Engine and the Fuel
Artificial Intelligence is often described as a powerful engine capable of transforming industries from agriculture to healthcare. But like any engine, it needs fuel to run. In AI, that fuel is data. The performance, accuracy, and reliability of any AI system
are directly tied to the quality of the data it is trained on. High-quality data leads to robust and useful applications, while poor-quality data results in faulty, inefficient, and sometimes harmful outcomes. According to some studies, up to 80% of the work in an AI project can be data preparation, highlighting its foundational importance. This makes the painstaking work of collecting, cleaning, and organising data the most critical step in building a successful AI ecosystem.
The 'Garbage In, Garbage Out' Problem
The principle of 'garbage in, garbage out' is especially true for AI. If a model is trained on flawed data, its outputs will be flawed. In the Indian context, 'bad data' can mean many things: incomplete records, inconsistent formats, outdated information, and poor labelling. A significant issue is the lack of well-annotated local datasets. For example, an AI model designed for crop disease detection will fail if it's only trained on images from other continents. Similarly, a financial AI that assesses loan applications will be ineffective if its data excludes the vast informal economy. This problem is compounded by India's immense diversity. Datasets that are not representative of the country's various languages, cultures, and socio-economic strata can lead to biased and inequitable AI.
Real-World Impact: Bias in Action
Algorithmic bias is not a theoretical problem; it has real-world consequences. In India, AI models have shown biases that reflect and amplify existing social inequalities. For instance, some AI-powered credit scoring systems have been found to systematically disadvantage rural applicants because their training data over-represents urban, salaried individuals. This means a farmer with a solid credit history in an informal system might be denied a loan that a city-dweller with less experience gets. Similarly, facial recognition systems have demonstrated higher error rates for individuals with darker skin tones because the models were predominantly trained on images of lighter-skinned people. AI models have even been shown to associate surnames with caste and profession, potentially leading to discrimination in hiring and credit assessment.
The Promise in Agriculture and Healthcare
The potential for AI in vital sectors like agriculture and healthcare is immense, but entirely dependent on data quality. In agriculture, AI can provide farmers with precise advice on everything from sowing times and irrigation schedules to pest management. These systems rely on analysing vast amounts of localised data, including satellite imagery, weather patterns, and soil health information. Without high-quality, granular data, these tools cannot provide the actionable insights needed to boost crop yields and improve food security. In healthcare, AI promises to help with diagnostics and treatment plans, but its effectiveness hinges on being trained with diverse Indian patient data. An AI trained only on Western medical data could misdiagnose conditions or suggest inappropriate treatments for an Indian population.
Building a Better Data Foundation
Recognising this critical need, India has launched several initiatives to build a robust data ecosystem. The IndiaAI Mission, approved in March 2024, has made enhancing data quality one of its central pillars. A key part of this strategy is the IndiaAI Datasets Platform, also known as AIKosh. This platform aims to create a unified repository of high-quality, non-personal datasets for researchers and startups, thereby lowering the barrier to innovation. By providing access to large-scale, diverse, and anonymised data, the government hopes to reduce bias and improve the reliability of AI applications across sectors. Initiatives like AIKosh, which invite both public and private entities to contribute datasets, are crucial for creating the shared infrastructure needed for responsible AI development.
















