Digital Colonialism

In the summer of 2021, a small company called Sama, operating out of an office tower in Nairobi, took on a contract that would help build one of the most famous products in the history of software. OpenAI needed a way to keep ChatGPT from spewing violence, hate, and sexual abuse at its users. To teach an AI to recognize toxic text, you first have to feed it examples of toxic text — thousands upon thousands of them, each tagged by a human who has read it and judged it. That reading and tagging is what the Sama workers did. For hours a day, Kenyans earning a take-home wage of between roughly $1.32 and $2 an hour scrolled through descriptions of child sexual abuse, bestiality, murder, suicide, and torture, labeling each one so that a model in California could learn to refuse it (Perrigo, TIME, 2023). Some of them later said the work left them unable to sleep, unable to be intimate with their partners, haunted by images they could not unsee. When the details came to light, the workers described the mental-health support they had been offered as thin to the point of uselessness.

Here is the part worth sitting with. ChatGPT is, among other things, a triumph of "responsible AI." Its safety filters are a selling point. And those filters were built on the exposed nerves of low-wage workers in the Global South, whose labor was invisible to every user typing a prompt into the clean white box. The value flowed north. The trauma stayed south. And the arrangement was not an aberration or a scandal at the edge of the industry — it was the industry working exactly as designed.

That pattern, repeated across data, labor, hardware, knowledge, and law, is what a growing body of scholarship calls digital colonialism: not conquest by armies but extraction through data flows, dependency through infrastructure, and marginalization through algorithms trained on someone else's reality. The term is deliberately provocative, and we will need to ask whether it is fair. But first we need to see the machine clearly. Digital colonialism is not one thing. It is five distinct dynamics — data extraction, ghost labor, infrastructure dependency, epistemic marginalization, and governance asymmetry — each of which operates on its own, and each of which quietly makes the others worse.

Five gears in one machine

It helps to picture the five dynamics as gears that turn independently but mesh. You could, in principle, fix any one of them without touching the others — a country could build data centers without gaining a seat at the governance table, or win labor protections for its data workers without changing whose epistemology its AI encodes. But left alone, the gears drive each other.

graph LR
  A[Data extraction] --> B[Ghost labor]
  B --> C[Infrastructure dependency]
  C --> D[Epistemic marginalization]
  D --> E[Governance asymmetry]
  E --> A
  C --> A
  E --> C

Data harvested from the periphery trains models that require compute the periphery does not own; the absence of local compute means local labor is contracted, not employed; models trained on foreign data encode foreign ways of knowing; and the rules that might correct any of this are written by the same governments and firms that benefit from leaving it in place. The reinforcement is the point. Understanding it means taking the gears one at a time.

The new extractivism

Old colonialism took gold, rubber, and minerals out of the ground. Digital colonialism takes data out of behavior.

Every quiz answered on a foreign learning platform, every symptom typed into a health app, every face scanned by a proctoring system, every search, swipe, and voice command generates a record that flows to servers in the Global North, where it becomes raw material for models that the people who generated it will never own. The theorists Nick Couldry and Ulises Mejías, who did more than anyone to name this dynamic, call it "data colonialism" precisely to insist that the appropriation of human life as data is a new frontier of extraction, continuous with the old one rather than merely analogous to it (The Costs of Connection, 2019; Data Grab, 2024).

"Data is the new oil" is the usual line, and it undersells the case. Oil is burned once. Data compounds. Each interaction produces more training data, which sharpens the model, which draws more users, which produce more data — a flywheel that hands whoever controls it a widening lead. And unlike oil, data carries culture inside it: language, categories, assumptions about what a normal patient or a good essay or a legitimate transaction looks like. When a model trained on one society's data is deployed in another, it does not arrive neutral. It arrives opinionated.

This is where the third of our supporting questions bites — why scaling without localization is a form of imposition even when nobody intends harm. The mechanism is not malice; it is market logic. A firm builds a product for the data it has and the users it understands, then scales it globally because scaling is what software is for. The marginal cost of serving one more country is near zero, so the incentive is to ship the existing model, not to rebuild it for local conditions. A learning platform flags a teaching method as ineffective because it underperforms against Western benchmarks. A content system penalizes a language it barely represents. None of this requires a colonial intent; it requires only that the cheapest path to a billion users runs through defaults set in Palo Alto. The imposition is structural, baked into the economics, which is exactly why it is so hard to argue away by pointing to good intentions.

The ghost workers

The Sama story is not unique; it is representative. Behind the appearance of automation stands a vast, deliberately hidden workforce. The anthropologists Mary Gray and Siddharth Suri named them "ghost workers" — the people performing the "last mile" of human judgment that AI cannot yet do for itself (Ghost Work, 2019).

The geography is consistent. Kenyans label images of streets for self-driving cars they will never ride in and moderate content for platforms headquartered continents away. Filipinos — by some industry estimates over a million of them at peak — annotate data and scrub feeds through firms in Manila and Cebu. Indians transcribe audio to train voice assistants tuned to American accents. Venezuelans, after their economy collapsed, flooded onto platforms like Remotasks and Scale AI because a dollar an hour in bolívars was worth having. The through-line is arbitrage: the same task that would trigger minimum-wage law, benefits, and mental-health obligations in San Francisco costs a fraction of that in Nairobi or Manila, with none of the obligations attached.

What structurally shields these workers from protection is the contract itself. They are classified not as employees but as independent contractors or as staff of a local outsourcing firm two or three steps removed from the AI company whose product they build. That distance is not incidental — it is the mechanism. It means the tech firm can truthfully say it does not employ them, the outsourcer can say it merely follows the client's specifications, and the worker can appeal to neither. Pay is piece-rate and unpredictable; an account can be deactivated without warning or explanation; tasks can dry up overnight when a contract moves to a cheaper country. For content moderators, the work is additionally traumatic in a way the industry has been slow to acknowledge: sustained exposure to the worst material humans produce, at a pace that leaves no room to recover between items.

There is, notably, a countercurrent here that matters for the rest of the chapter. Kenyan moderators have organized, sued, and forced the conditions into open court — legal actions against Meta and Sama that have tested, for the first time, whether a Global North platform can be held liable in a Global South jurisdiction for the labor that trains and cleans it. Those cases are unresolved as of this writing, and their outcome is one of the clearer near-term tests of whether ghost labor stays invisible or becomes governable.

The infrastructure dependency

Africa is home to roughly 18% of the world's people and, by widely cited industry estimates, less than 1% of its data-center capacity. Sit with the arithmetic. It is not a rounding error or a temporary lag; a gap that large is a structural feature, and it dictates the terms on which an entire continent can participate in the AI era.

The reason infrastructure matters so much is that modern AI is not software you can simply download. Training and running large models demands enormous, specialized, power-hungry compute concentrated in physical buildings. A country without that capacity cannot host advanced AI workloads locally. It must rent them — from Amazon Web Services, Microsoft Azure, or Google Cloud, all of which run on machines sitting in the Global North, governed by Northern law, priced in Northern currencies, and switchable off by Northern decision. That is not a commercial detail. It is a dependency with a political edge: what happens to a Nairobi hospital's diagnostic system, or a Jakarta ministry's records, if pricing shifts, sanctions bite, or a provider decides a market is no longer worth serving?

The fourth supporting question asks why this dependency is circular, and the answer is almost cruel in its logic. To escape reliance on foreign cloud, a country needs its own data centers. To build data centers, it needs cheap reliable power, deep capital, technical talent, and stable regulation. But the countries most dependent on foreign cloud are precisely the ones short on one or more of those inputs — and their dependence keeps the domestic market too small and too risky to attract the investment that would build the alternative. Dependency suppresses the conditions for autonomy, and the absence of autonomy deepens dependency. What breaks the cycle is rarely the market on its own; it is deliberate, patient, usually public or multilateral investment that treats compute as infrastructure — closer to roads or the grid than to a consumer product — and is willing to accept returns measured in strategic capacity rather than quarterly profit.

China read this situation early and moved. Its Belt and Road Initiative acquired a digital limb — the "Digital Silk Road" — under which firms like Huawei have laid submarine cables, built data centers, rolled out 4G and 5G networks, and installed "safe city" surveillance systems across Africa, South and Southeast Asia, and Latin America, often on generous financing terms that Western firms would not match.

The fifth question is the sharp one: does this give recipient countries genuine autonomy, or merely a different master? The honest answer is that it substitutes one dependency for another with a different risk profile. Chinese infrastructure comes with its own strings — debt owed to Chinese lenders, systems that lock in Chinese vendors and standards, surveillance architectures that serve incumbent governments and, potentially, Beijing's intelligence interests. Western cloud dependency and Chinese infrastructure dependency are not the same thing, and pretending they are obscures the choice. Western dependency is a rental relationship: you get frontier capability and no ownership, exposed to the provider's commercial and geopolitical whims. Chinese dependency is more like a mortgage: you get physical assets on your own soil, and a long-term relationship of debt and technical lock-in to the lender. A country genuinely gains something under the Chinese model that pure cloud rental never delivers — hardware it can point to, some transfer of operational skill. Whether that adds up to more autonomy or less is one of the real open questions in this field, and the next section returns to how little we actually know about it.

The epistemic marginalization

The fourth gear is the subtlest, because it operates on knowledge itself. An AI system trained overwhelmingly on one civilization's data does not merely reflect that civilization's world — it starts to define what counts as the world.

Medicine shows the mechanism at its most concrete and most dangerous. A diagnostic model trained on European or North American patients learns the disease patterns, the symptom presentations, and the demographic baselines of those populations. Deployed in sub-Saharan Africa or South Asia, it may miss conditions rare in its training data and misread symptoms that present differently across bodies. This is not speculation; the adjacent evidence in medical devices is stark. Pulse oximeters, calibrated largely on lighter skin, were found to be roughly three times more likely to miss dangerously low oxygen in Black patients than in white ones (Sjoding et al., New England Journal of Medicine, 2020) — a bias that ran quietly through hospitals for years and shaped who got treated during the COVID-19 pandemic. Dermatology AI carries the same flaw: skin-cancer classifiers were trained on datasets containing very few images of dark skin, so their impressive published accuracy simply does not describe how they perform on the patients most of the world's population actually has. Even when a model is technically excellent, deployment reveals the gap: Google's diabetic-retinopathy system, superb in the lab, stumbled in Thai clinics where lighting, connectivity, and workflow did not match the conditions it was built for (Beede et al., 2020). A system built with global deployment in mind but local data at its core gets treated as universal precisely where it is most parochial.

For students and scholars, the marginalization is quieter but pervasive. The overwhelming majority of large language models are optimized for English and a handful of other high-resource languages; performance degrades sharply across most African and many South Asian languages, which means the tool that answers a London undergraduate's question fluently gives a Lagos or Addis Ababa student something halting, thin, or wrong. Citation and search systems privilege English-language journals, so a scholar publishing in Kiswahili or Amharic is not merely under-read — they are, to the machine, nearly invisible. And the categories run deeper than language. A platform organized around Western academic disciplines may have no slot for Indigenous ecological knowledge, oral historiography, or non-Western medical and philosophical traditions; a student working in those frameworks can find that the system cannot recognize, let alone assess, what they know. The knowledge is not judged inferior. It is not registered at all — which is worse, because invisibility does not even announce itself.

The governance asymmetry

The fifth gear is the one that could, in principle, adjust all the others — and it is jammed in the same direction as the rest.

The rules that will govern AI are being written overwhelmingly in the Global North. The European Union's AI Act, the executive orders and agency actions of the United States, and the technical standards emerging from bodies like ISO and IEEE are shaped by institutions from wealthy, technologically dominant states. The empirical base is skewed the same way: by the most-cited estimates, well under 1% of global AI research funding reaches the Global South, so even the studies that inform governance describe the North's concerns using the North's data.

The EU's reach is amplified by what the legal scholar Anu Bradford named the Brussels Effect (2020): because it is cheaper for a multinational to apply one strict European standard everywhere than to maintain a patchwork, European rules become de facto global rules — extending to users in Jakarta or São Paulo who had no vote in Brussels and whose priorities may differ sharply. Regulation designed for a rich, data-saturated, litigious market gets exported to contexts it was never meant to fit.

The eighth question asks what leverage the Global South actually has, and the answer is real but conditional. When a middle-income country drafts assertive AI rules, it runs into predictable friction: platforms hint they may withhold services or investment; international bodies press for "harmonization," which usually means adopting Northern templates wholesale; and the country often lacks the trained policy experts and the access to proprietary systems needed to regulate on its own terms. Leverage becomes effective under a narrow set of conditions — chiefly market size and the willingness to coordinate. India can extract concessions that Malawi cannot, because losing the Indian market is a board-level event. The Brussels Effect itself is proof of the principle: a market large enough to be indispensable sets terms. That is why the most promising Global South strategy is bloc-building — the African Union's continental AI strategy, regional coordination in Latin America, and India's repeated positioning of itself as a voice for the "Global South" in AI forums are all attempts to manufacture, through aggregation, the leverage that no single mid-sized economy holds alone.

Is it still setting, or has it set?

A fair reader should want to know where we are in the arc. Is digital colonialism an emergent dynamic still consolidating — and therefore reversible — or has it already hardened into self-reinforcing structural dependency?

The distinction is not academic. An emergent dynamic can be redirected by policy while the concrete is wet; a hardened one requires breaking existing structures, which is far costlier and rarely happens without crisis. What would distinguish the stages? In the emergent stage, we would expect the dependencies to be shallow and switchable, alternative providers to be viable, and local capacity to be growing fast enough to catch up. In the hardened stage, we would expect lock-in — switching costs so high that dependence persists even when countries want out — and a widening, not narrowing, capacity gap.

The evidence tilts toward hardening, though not uniformly. The compute gap is widening, not closing: frontier training runs demand capital and energy at a scale that puts them further out of reach each year, so a continent at under 1% of capacity is falling behind the frontier even as it adds data centers in absolute terms. Cloud lock-in is deep and deliberate; migrating a national health system off AWS is not a weekend project. Foundation models trained today entrench English and Western data as the substrate on which everything downstream is built. Against that, the labor gear is arguably still emergent — the Kenyan lawsuits show the rules are genuinely unsettled — and, as we'll see, indigenous capacity is appearing faster than the pessimistic reading predicts. The most defensible summary: infrastructure and epistemic dependency are hardening toward structural permanence, while labor and governance remain contested enough that the outcome is not yet written.

Cracks of autonomy

If the structure is hardening, the tenth question matters most: what do the early attempts at indigenous capacity actually show about realistic timelines?

India's BHASHINI is the most cited example — a national initiative to build AI translation and speech tools across the country's constitutionally recognized languages, so that digital services can reach citizens in their own tongue rather than in English. It is ambitious, state-backed, and explicitly framed as digital public infrastructure rather than a product. In Africa, the picture is grassroots as much as governmental: Masakhane, a distributed, largely volunteer community of researchers, has done more than any corporation to build natural-language processing for African languages, and startups like South Africa's Lelapa AI have released small language models purpose-built for African languages rather than bolted onto English foundations. A scatter of African AI research centers and national strategies rounds out the trend.

What these examples suggest is sobering but not defeatist. They prove that meaningful local capacity is possible, and that it grows fastest where three conditions coincide: sustained public funding, a talent base retained rather than drained abroad, and a deliberate choice to treat the tools as public infrastructure. But they also reveal the ceiling. None of these efforts competes at the compute-hungry frontier of the largest models, and none is close to doing so; they carve out sovereignty at the application and language layer while remaining dependent on foreign compute and foreign foundation models underneath. Realistic autonomy, on this evidence, is a decade-scale project measured not in matching OpenAI but in owning the layers closest to citizens — language, data, services — while the base layer stays, for now, rented.

How well do we actually know this?

This book insists on separating what we know from what we suspect, and digital colonialism is a domain where that discipline matters, because the moral charge of the word "colonialism" can outrun the evidence.

Some of the harms are exceptionally well documented. Ghost-labor conditions rest on firsthand investigative reporting, worker testimony, and court filings — the Sama wage figures and working conditions are as solid as this kind of evidence gets. The infrastructure gap is measurable and consistently reported across independent industry analyses. The performance gaps of Western-trained models in non-Western contexts are demonstrated in peer-reviewed medical literature and reproducible NLP benchmarks. On these, confidence is warranted.

Other claims are softer, and honesty requires saying so. "Data extraction without compensation" is well established as a description of how the business model works, but quantifying the value extracted from a given population — putting a number on what is owed — remains largely estimative. And the fifteenth question points to a genuine hole: we do not actually know whether China's Belt and Road digital investment leaves recipient countries better or worse off than Western cloud dependence. The comparison is asserted far more often than it is measured. A fair test would track matched countries over time across cost, reliability, skills transferred, surveillance risk, and exit options under both models — and that study, as far as the public record shows, has not been done. Anyone who tells you confidently which dependency is better is reasoning from ideology, not data.

Which brings us to the framing itself. Is "digital colonialism" analytically useful, or does it inflate an analogy in ways that mislead? The strongest objection is real: historical colonialism meant territorial conquest, direct political rule, chattel slavery, and violence on a scale that data extraction, however exploitative, does not match — and Global South countries mostly adopt these technologies voluntarily, often gaining genuine benefits, which is not how conquest works. Stretch the word too far and it flattens those differences, and can even insult the memory of what colonialism actually was. The strongest defense is equally real: the term correctly names a structural pattern — value flowing from periphery to center, dependency engineered and maintained, the periphery governed by rules it did not write — that milder language ("digital divide," "AI gap") obscures by making it sound like a natural lag rather than a produced relationship. The most useful position holds both: digital colonialism is a precise description of a structure and a loose description of an event. It illuminates the political economy and misleads if taken as a literal historical equation. Getting this right is not pedantry — it determines the policy response. If the problem is a "gap," you fund catch-up. If it is a produced dependency, you have to change who owns what and who writes the rules. The framing chosen decides which solution even comes into view.

What is owed, and what would make it real

The normative questions follow directly. Should Global South data be treated as a compensable resource? The case is strong: if data is the raw material of a trillion-dollar industry and it is generated by specific people and institutions, treating it as free-to-take is a choice, not a law of nature. The hard part is operationalizing it. The proposals on the table — data trusts that hold and license data collectively on a community's behalf, collective-ownership and data-sovereignty regimes that keep data under local legal control, and revenue-sharing or local-reinvestment requirements on firms that profit from a region's data — all rest on the same principle: those who generate the data should hold rights over it and share in its value. Each faces genuine technical and legal friction, and none is yet proven at scale, but the principle is coherent and the experiments are underway.

The obligations of Global North firms follow the same logic, and the enforcement question is where good intentions usually die. Voluntary corporate commitments to "responsible AI" have, on the evidence of the Sama contract, not protected the people at the bottom of the supply chain. What could make obligations real is a shift from voluntary to binding: supply-chain due-diligence law that holds a firm liable for labor conditions several contracting steps removed from it (the direction the Kenyan lawsuits are testing), procurement rules that condition market access on local investment and data-sovereignty compliance, and — crucially — jurisdiction, so that a worker harmed in Nairobi can seek remedy against the firm that ultimately profited. Enforcement, not exhortation, is the variable.

Finally, the governance bodies themselves. Nominal inclusion — a seat at a forum, a name on a communiqué — is not the same as genuine voice, and the thirteenth question asks what would close that gap. Three reforms recur across the serious proposals: representation with actual decision-making authority rather than observer status in standard-setting bodies; funding mechanisms that build regulatory and research capacity in underrepresented regions, so that participation is informed rather than symbolic; and decision rights structured so that the Global South is not permanently outvoted by the wealthy states that also happen to host the industry. The UN's recent moves toward a global AI advisory role and a Global Digital Compact gesture in this direction, but they remain advisory, and advice without authority is exactly the nominal inclusion the reform is meant to replace.

Summary

Digital colonialism is not a metaphor for a single injustice but a name for five distinct, mutually reinforcing dynamics.

  1. Data extraction. The behavior of Global South users trains and enriches models owned in the North, in an exchange that returns little to those who generate the data — a structural imposition that market logic produces without anyone intending harm.
  2. Ghost labor. The human work that makes AI function — labeling, moderation, transcription — is performed by low-wage, contractually distanced workers in Kenya, the Philippines, India, and elsewhere, shielded from protection by the very outsourcing structures that hide them.
  3. Infrastructure dependency. With Africa holding under 1% of global data-center capacity against 18% of the world's population, most of the Global South must rent compute from Northern providers or borrow it from China — a circular trap where dependency suppresses the conditions for autonomy.
  4. Epistemic marginalization. Models trained on Western data carry Western assumptions into medicine, language, and scholarship, producing real diagnostic failures and rendering non-Western knowledge invisible when the tools are treated as universal.
  5. Governance asymmetry. The rules are written where the industry lives; the Global South is governed by frameworks it did not design, with leverage available mainly to those large enough — or coordinated enough — to be indispensable.

The evidence for these harms is strong on labor, infrastructure, and performance gaps, and thinner on the precise value extracted and on whether Chinese infrastructure serves recipients better than Western cloud — a genuinely open question. The "colonialism" framing is precise about structure and loose about history, and which reading you adopt determines whether the fix looks like charity (close the gap) or like justice (change who owns what and who writes the rules). Early efforts — India's BHASHINI, Masakhane, Lelapa AI, a scatter of national strategies — prove that indigenous capacity is possible at the application and language layers, on a decade-scale timeline, while the compute frontier stays out of reach. Left to the market, these dynamics harden. Redirecting them requires the deliberate choice to treat data as compensable, obligations as enforceable, and governance as something the governed help design.

Sources

Last updated: 2026-08-16

V2 (in progress) Previous: V1