Cultural Homogenization vs. Diversity
Keahi speaks ʻŌlelo Hawaiʻi—Hawaiian—fluently. His grandmother taught him. She had learned from her grandmother, who learned it before the language was banned in schools, before it nearly died.
For most of the twentieth century, Hawaiian was on the edge of extinction. After the 1896 law that made English the sole language of instruction, transmission collapsed; by the 1980s fewer than a thousand native speakers remained, most of them elderly. Then came the revival—immersion preschools called Pūnana Leo, university programs, a deliberate act of cultural reclamation. Today tens of thousands speak it again. Hawaiian is one of the rare cases where a language walked back from the grave.
Keahi teaches at one of those immersion schools. In 2024 he tried an AI translation tool to speed up making classroom materials. He typed in Hawaiian phrases and asked for English. The results were not just clumsy—they were culturally incoherent. The model rendered aloha as "hello" or "goodbye." Technically defensible. Profoundly wrong. Aloha is breath and presence, compassion, the mutual regard between people; it is a philosophy compressed into five letters. The tool had been trained overwhelmingly on English, with a thin scraping of Hawaiian websites layered on top. It didn't understand Hawaiian. It pattern-matched Hawaiian into an English-shaped mold, and in doing so it quietly hollowed the language out.
He stopped using it. But the episode points at something larger than one classroom. How many learners are studying a minority language through tools that teach them a version of it filtered through English—the shape of the language without its interior? Multiply that across thousands of tongues, and you have a machine that can erase linguistic diversity while advertising that it preserves it.
How endangered, and how uneven
Start with the raw arithmetic, because it frames everything else. There are roughly 7,000 living languages in the world (Ethnologue counts about 7,160). Somewhere around 40 percent of them are endangered—commonly defined as having too few speakers, and too little intergenerational transmission, to survive the century. UNESCO's long-standing estimate is that a language dies roughly every two weeks. About 3,000 languages have fewer than 1,000 speakers each; many of those will be gone within a generation or two.
Now lay AI coverage on top of that map, and the unevenness is staggering. The most cited attempt to measure it, Joshi and colleagues' "The State and Fate of Linguistic Diversity in the NLP World" (Joshi et al., 2020), sorts languages into six tiers by how much digital data and tooling exist for them. At the top—tier five—sit a handful of "winners": English, and a small club including Mandarin, Spanish, French, German, Japanese, Arabic. At the bottom—tier zero, the "left-behinds"—sit roughly 2,200 languages, around 88 percent of all languages, with almost no digital resources at all. English alone accounts for something close to half of the text in the web-scraped corpora that train large models, while it is the native language of under 5 percent of humanity.
The correlation between speaker population and AI quality is real but—and this matters—imperfect, and the imperfection is revealing. If it were purely about headcount, the largest languages would all be well served. They aren't. Hindi has over 600 million speakers and is still comparatively under-resourced; Bengali, Swahili, Javanese, and Yoruba all support enormous populations while lagging far behind languages with a fraction of their speakers. What actually predicts good AI coverage is not how many people speak a language but how much of that language sits online in machine-readable text—and that, in turn, tracks wealth, colonial history, and which communities got early internet infrastructure. AI coverage is less a map of the world's speakers than a map of the world's servers.
Why the models think in English
The mechanism is not conspiracy; it is gravity. A large language model learns the statistical shape of whatever text it eats. The text available to eat is overwhelmingly English and a few other high-resource languages, because that is what has been written, digitized, and made scrapeable over three decades of a mostly English-language internet. So the model's internal representations are built in an English-shaped space, and everything else is bent to fit.
This persists even when a company advertises multilingual support. Meta's No Language Left Behind project claimed 200 languages; models like BLOOM and Cohere's Aya push past 100. But "supported" and "supported well" are different claims. For low-resource languages, benchmark accuracy drops sharply, hallucination rises, and outputs drift toward the grammar and assumptions of the dominant training languages. A student in Odisha analyzing a paper in Odia will often find the leading models simply let them down—defaulting to English patterns, producing translations that are functional but culturally off-key.
There is also a cost mechanism that most users never see. Large models break text into "tokens," and they are billed and rate-limited per token. Because their tokenizers are optimized for English, the same sentence in an under-resourced language can require several times as many tokens—in some scripts up to fifteen times as many (Petrov et al., 2023, "Language Model Tokenizers Introduce Unfairness Between Languages"). That means speakers of those languages pay more per query, wait longer, and hit context limits faster, for a worse result. The disadvantage is baked into the plumbing, not just the output.
What flattens, and what is lost
Languages are not interchangeable label-sets for one shared reality. They encode different ontologies—different ways of carving up kinship, time, obligation, and the natural world. When a model translates, it must map concepts from one system onto another, and often they don't map cleanly. It faces a choice: translate literally, producing something opaque, or translate idiomatically, producing something readable that has quietly lost the original. Most systems choose readable. The distinctive features get sanded away to fit the target language's frame.
The Hawaiian ʻohana is a clean illustration. Rendered as "family," it isn't wrong—but ʻohana reaches past the nuclear household to extended and chosen kin, and carries binding obligations of mutual care that the English word doesn't hold. Flatten ʻohana to "family" a million times across a million translations and you don't just lose a word; you erode a relational worldview. The same happens to Portuguese saudade, Japanese mono no aware, the fine-grained Sámi and Inuit vocabularies for snow and sea ice. Each is a structure of meaning that resists clean substitution.
This is why the death of a language is not merely a linguistic event. When the last fluent speaker of a language dies, what disappears is not recoverable from a dictionary. Much of what a language carries—ecological knowledge encoded in plant and animal names, oral histories, navigational and medicinal traditions, ways of reasoning about relationship and time—lives only in use, in the heads of speakers, and was never written down. A dictionary preserves the bones; the living body of knowledge is gone. That is why the chapter treats each loss as a subtraction from the total stock of human knowledge, not a cosmetic change in the world's vocabulary. A Hawaiian proverb states the stakes without decoration: "I ka ʻōlelo ke ola; i ka ʻōlelo ka make"—in language there is life; in language there is death.
The pressure that pushes languages downhill
For the speaker, the consequence of a bad tool is concrete and cumulative. If your language's AI stumbles, then education, work, government forms, search, and online life all run more smoothly in English—so you switch to English for anything practical. Each switch withdraws a little more of the everyday utility that keeps a language alive. Younger people, watching their elders' language handle their devices poorly, quietly conclude it isn't worth learning well. Transmission is where languages live or die, and transmission is exactly what erodes when the practical payoff drains away. AI doesn't outlaw a language. It makes it inconvenient—and in a globalized economy where utility governs survival, inconvenience is often enough.
This is the honest core of the epistemic question: how much of today's endangerment can we actually pin on AI? The candid answer is that AI is an accelerant, not the arsonist. Languages were already dying at speed from urbanization, state-language schooling, economic migration, and the pull of a globalized labor market—forces that predate the technology by a century. AI's contribution is that it hardens the incline. It adds one more domain—the digital one, now central to daily life—where the dominant languages work and the small ones don't, and it does so at planetary scale and machine speed. The trajectory is not new. AI shortens the timeline.
The content monoculture
The homogenizing pressure runs past language into culture itself, through two distinct mechanisms worth keeping separate.
The first is generation. AI models trained on a web that skews Western and English produce outputs that inherit that skew. Ask for an image and the default aesthetic is Western; ask for a story and it tends toward Western narrative arcs; ask for music and it draws on Western tonal traditions. No one at the company decided this. It falls out of the data. But a teenager in rural Indonesia growing up on AI-generated video, images, and stories is absorbing a stream that is subtly Western in its assumptions—and against that steady current, gamelan, wayang puppetry, and Javanese storytelling start to feel exotic and old-fashioned in their own home. This is not colonialism in the old sense; no one is imposing anything. It is passive erosion through the sheer default of the tool, which is in some ways harder to resist precisely because there is no villain to point at.
The second mechanism is curation. Recommendation systems—YouTube, TikTok, Spotify, Netflix—optimize for engagement, which means showing people more of what already holds them. At the scale of billions of users, that individual optimization produces collective convergence: tastes align, the same songs and formats and memes spread everywhere, and genuinely local or niche production gets deprioritized because it can't match the engagement metrics of globally optimized content. An Icelandic band rooted in traditional folk forms may make something genuinely distinctive—and the algorithm, having no engagement history for it, rarely surfaces it beyond a tiny existing audience. The band adapts toward global trends and loses what made it interesting, or stays obscure. The paradox is sharp: AI grants everyone access to the entire cultural output of the species, then filters that access down to a narrow, engagement-optimal slice. You can hear anything. You mostly hear what the model predicts you'll finish.
How much of this convergence is demonstrated rather than merely theorized? Here we should be careful. The mechanics of engagement optimization are well understood, and there is solid evidence of filter bubbles and of hit concentration in streaming markets. But the strong claim—that global culture is measurably collapsing into one form—is still partly an extrapolation from those mechanics, and it runs against real counter-evidence: the same platforms carried K-pop, Afrobeats, reggaetón, and Nigerian cinema to global audiences that gatekept media never would have. The convergence pressure is real; whether it dominates the diversifying pressure is genuinely unsettled, and honesty requires holding both at once.
What communities are building instead
The most encouraging response is not coming from the large platforms. It is coming from communities building their own tools, on their own data, on their own terms.
Te Hiku Media in New Zealand is the standout. Working with fluent speakers and, later, with NVIDIA, they built automatic speech recognition for te reo Māori that reached around 92 percent accuracy—remarkable for a low-resource language. Crucially, they did it under a licence that keeps the data and the model under Māori control, refusing to hand the community's recordings to outside firms to monetize. Masakhane, a grassroots network of researchers and native speakers across Africa, has built translation and NLP tools for dozens of African languages rather than waiting for Silicon Valley to get around to them. Researchers at Dartmouth have partnered with tribes to document and model Native American languages. The pattern across the successes is consistent: fluent elders and linguists in the loop, authentic vetted texts rather than scraped fragments, community ownership of the data, and tools designed to support human teaching rather than replace it.
This is the substance of the argument for community-led, data-sovereign AI over patience with the big platforms. A commercial model optimizes for the largest markets and the cleanest available data; it has no incentive to serve a language of 4,000 speakers, and every incentive to scrape whatever it can without consent. A community model can encode the dialect the community actually wants preserved, can refuse to launder colonial-era distortions back into the language, and can treat fidelity rather than engagement as the target. The Indigenous data sovereignty movement—organized around principles like CARE (Collective benefit, Authority to control, Responsibility, Ethics)—frames this as a right: communities deciding how their heritage is digitized and used, not a courtesy extended to them afterward.
The window, and who gets to close it
Two hard questions remain, and they point the same direction. First, the timing. For a language with fewer than 1,000 speakers, the window is not measured in decades—it closes when the last fluent generation does, often fifteen to thirty years. AI-assisted documentation has to reach them while they can still record, correct, and teach; after that, a model can only mimic what was captured, not restore what wasn't. Second, the allocation. The communities most in need of this support—small, under-resourced, digitally thin—are precisely the ones market forces will never reach, because there is no business case for a language that cannot fund an API bill. Left to the market, the gap widens: the well-resourced languages get better tools every quarter while the endangered ones get nothing, and the relative distance grows even as absolute coverage of low-resource languages slowly improves.
That mismatch is what grounds the normative claim. If a handful of companies now shape, through their default outputs, how billions of people write, learn, and consume culture, then the scale of that influence arguably creates an obligation that market incentives alone won't discharge—to fund coverage for languages that will never pay for themselves, and to design curation systems that deliberately preserve exposure to cultural difference rather than converging on whatever maximizes watch time. The obvious objection is who decides what counts as adequate diversity—a question with no neutral answer, and one that a San Francisco product team is poorly placed to settle. Which is exactly why the answer keeps returning to the communities themselves: the least bad arbiter of what a culture needs to survive is the culture, holding its own data.
The genuinely open question—the one nobody can yet answer—is whether the tools now being built for minority languages are helping or subtly harming. A model trained on thin or unrepresentative data may encode a distorted dialect, a colonial-era register, or plain errors, and then teach that flawed version to a generation of learners who have no fluent elder to correct it. The language would appear to be reviving on the dashboards while quietly mutating into something its ancestors wouldn't recognize. Detecting that harm requires exactly what the endangered communities most lack: enough fluent speakers to audit the output. The measurement problem and the endangerment problem are the same problem.
Summary
-
The scale is large and the coverage is lopsided. Around 7,000 languages exist; roughly 40 percent are endangered, and about 3,000 have fewer than 1,000 speakers. Yet close to 88 percent of languages have almost no digital resources at all (Joshi et al., 2020). AI coverage tracks not speaker numbers but the amount of a language available online—so it maps the world's servers, not the world's speakers.
-
English-centrism is structural, and persists under multilingual branding. Models learn the shape of their training data, which is overwhelmingly English. Even "200-language" systems perform far worse on low-resource languages, and tokenizer design makes those languages literally more expensive and slower to use (Petrov et al., 2023).
-
Translation trades fidelity for readability. Idiomatic translation flattens concepts like ʻohana and aloha that have no clean English equivalent, gradually stripping languages of the worldviews they encode.
-
AI accelerates rather than initiates language loss. Urbanization, globalization, and economic marginalization were already killing languages; AI hardens the incline by adding one more high-stakes domain where small languages don't work well.
-
Cultural homogenization runs through two mechanisms—generation defaulting to Western aesthetics, and engagement-optimized curation converging tastes—but the strong "global monoculture" claim is still partly extrapolation, with real counter-evidence from globally viral non-Western culture.
-
Community-led, data-sovereign projects are the most viable path. Te Hiku Media, Masakhane, and Dartmouth's work show what fluent-speaker involvement and community ownership make possible. The window for the smallest languages is one generation wide, market forces will never reach them, and whether current tools truly help or subtly distort remains genuinely unknown—because auditing the harm needs the very speakers whose scarcity is the crisis.
Sources
- How many languages are endangered? | Ethnologue
- The State and Fate of Linguistic Diversity and Inclusion in the NLP World (Joshi et al., 2020) | arXiv
- Language Model Tokenizers Introduce Unfairness Between Languages (Petrov et al., 2023) | arXiv
- Why generative AI needs to be trained on more languages | World Economic Forum
- How AI Threatens Linguistic Diversity | China Daily
- Te Hiku Media and AI for Te Reo Māori | NVIDIA Blog
- Masakhane: AI for African Languages
- Language preservation efforts get an AI boost | Dartmouth
- How Multilingual AI Can Protect Language | Tech Policy Press
- Indigenous Data Sovereignty (CARE Principles) | Global Indigenous Data Alliance
Last updated: 2026-08-04
V2 (in progress) Previous: V1