The Illusion of Objective Data
We tend to think of data as a neutral, objective reflection of the world. In reality, data is never truly raw. It is collected, cleaned, and categorized by people, and trained on historical information that is often riddled with pre-existing societal
biases. This means that a database, no matter how many terabytes it holds, is not a perfect mirror of society but a portrait painted with the biases of its creators and the history it represents. From a sociological perspective, this isn't just a technical glitch; it's a structural reflection of social inequalities embedded in our institutions and cultural norms. When this flawed data is used to train artificial intelligence systems, the algorithms learn and can even amplify these underlying biases, creating a cycle of automated inequality.
When Good Tech Delivers Bad Outcomes
The consequences of this social narrowness are not theoretical. In fields like finance and hiring, algorithms designed for efficiency can become engines of exclusion. An AI recruitment tool trained on a company's past hiring decisions might learn to favour male candidates if the company has historically hired more men, perpetuating gender inequality. Similarly, a loan application system might use a postal code as a proxy for risk, unintentionally penalizing applicants from lower-income or minority neighbourhoods, a practice known as digital redlining. These systems don't need to be explicitly programmed to discriminate; the bias is inherited from the patterns in the data they analyze. The result is that people can be denied jobs or credit not because of their own merits, but because the algorithm has flagged them based on flawed, socially narrow data.
The Indian Context: A Magnified Challenge
In a country as diverse and socially complex as India, the risk of technically rich but socially narrow databases is particularly high. India’s deep-seated structural inequalities along the lines of caste, religion, gender, and region can be easily encoded into AI systems. For example, with a large portion of the population lacking internet access, especially women and rural communities, entire groups can be missing or misrepresented in datasets. This leads to skewed outcomes where systems are built around the problems of the affluent, like cardiac disease, while neglecting issues like tuberculosis that affect the poor. Furthermore, criminal databases that are used to train predictive policing tools may over-represent minority communities, leading to biased surveillance and arrests. Even generative AI can reflect these biases; chatbots asked to name Indian professors often return a list of dominant-caste surnames, highlighting how unequal representation in data leads to biased outcomes.
Aadhaar and the Human Cost of Exclusion
India's Aadhaar system, one of the world's largest biometric databases, provides a stark case study. While designed to streamline access to welfare, its implementation has highlighted the human cost of data-related failures. For manual labourers with worn-out fingerprints or elderly citizens whose biometrics have changed, authentication failures can mean being cut off from essential food rations or pensions. There have been multiple reports of deaths linked to starvation after families were denied benefits due to issues with linking their ration cards to Aadhaar. Errors in the demographic data, combined with the difficulty in correcting them, create endless hurdles for the poor. These issues demonstrate that a system can be a technical triumph in terms of scale, yet socially narrow enough to exclude the most vulnerable people it is intended to help.
Building Broader, More Inclusive Systems
Addressing this problem requires more than just better code; it requires a fundamental shift in how we approach technology. It starts with acknowledging that data is not neutral and building diverse teams of technologists, social scientists, and ethicists who can identify potential biases before they are encoded into a system. Companies and governments must move beyond a purely quantitative approach and incorporate qualitative, human-centred research to understand the social contexts in which their tools will operate. This means actively seeking out data from underrepresented groups and conducting rigorous audits for fairness, not just accuracy. Without a focus on transparency and accountability, even the most advanced databases will continue to be technically rich but socially blind, perpetuating the very problems we hope they will solve.
















