With the global market for AI data preparation expected to reach $1.2B by the end of 2023, India is emerging as a hub for data labelling and annotation work
Kumaramputhur is a tiny village some 45 km northwest of Palakkad in Kerala, home to some 3,500 families and probably not much bigger than an average Bengaluru suburb. Tweets: @kris_sg and @pankajontech Tweets: Kris Gopalakrishnan / @kris_sg : How India's data labellers are powering the global AI race | These are the jobs of the 21st century; massaging and curating data, removing offensive content, tagging and labeling etc. https://factordaily.com/... Pankaj Mishra / @pankajontech : How data labellers in Kumaramputhur, near Palakkad in Kerala, are powering the global AI race. Mujeeb Kolasseri, a high-school dropout, leads a team of over 200 employees working on AI solutions for clients across the world. By my colleague @murali_anand http://factordaily.com/...
Context & Ripple Effects
A month after FactorDaily reported that India's own researchers were starved of high-quality local datasets despite the country's data abundance (the local-dataset gap), the same outlet profiles where the export side of that business actually lives: Kumaramputhur, a Kerala village of some 3,500 families, where Mujeeb Kolasseri runs a team of over 200 labellers tagging and curating data for international AI clients.
The timing matters because market research put the global AI data-preparation market on track for $1.2B by end of 2023 — and the coverage since shows India capturing successive layers of it, from Karya's worker-owned training-data nonprofit in 2023 to Uber pushing labelling microtasks to its Indian drivers and the $250B IT services industry repositioning around data cleanup.
First-order effects
- Over 200 workers in Kumaramputhur hold AI-economy jobs — tagging, labelling, content moderation — without leaving a village 45 km from Palakkad, decoupling this work from metro BPO hubs like Bengaluru.
- Global AI clients gain a low-cost, English-capable annotation workforce outside the usual outsourcing centers, feeding directly into the $1.2B data-prep market.
Second-order effects
- New organizational models emerge to contest who captures the value: Karya's nonprofit structure redirects all profit to workers who keep ownership of their data, setting an explicit counter-model to village-scale vendor operations.
- Distribution platforms move in — Uber's AI Solutions arm routes photo-tagging microtasks to its existing Indian driver base, turning a rideshare app into a labelling workforce channel.
Third-order effects
- Data preparation is consolidating into a distinct Indian industry layer: with GenAI disruption pushing the $250B IT sector toward exactly this preparatory work, annotation shifts from informal village vendors toward platform-distributed gig labor and structured nonprofits.
- If the pattern holds, India's role in the AI supply chain extends beyond engineering talent — Bain projects a 1M-worker AI talent gap by 2027 — into being the default human-in-the-loop layer for global model training, raising questions about wages and data ownership that models like Karya are already testing.
The trend: AI's human-labour layer is industrializing in India, spreading from village vendor shops to gig platforms and worker-owned nonprofits as data preparation becomes core AI infrastructure.