In the earlier debates and subsequent essays that would form his seminal work, Decolonizing the Mind, Ngũgĩ argued that language was not merely a tool of communication, but also a vessel of culture, identity and memory. Ngũgĩ asserted that the colonial project had not ended with swapped flags, independence speeches or the ascent of African leaders, but that it had endured into how Africans thought, spoke and educated themselves. This radical conviction led him to cease writing in English, returning to his native Gikuyu, and to urge other African writers to divest colonial languages in favour of their native ones. Now, however, in an era shaped by the dominance of AI—with machines learning to speak and think in the world’s languages—Ngũgĩ’s insights on the linguistic hierarchies of the colonial project resonate with renewed force. And while Ngũgĩ’s call then was to reclaim language by thinking and writing in native tongues, today, the struggle is about embedding them into the data and architecture of the rapidly expanding global AI engine.
In a recent paper titled ‘Afrobench: How Good are Large Language Models on African Languages?’, a team of researchers put twelve large language models to test on 15 unique tasks across some 64 African languages. The results of the test revealed that when compared to their efficiency in high resource languages, popular AI models stumbled when prompted in the select African languages. Equally, other studies have consistently shown how large language models consistently distort or misattribute cultural references, or struggle entirely with local linguistic nuance. This inevitably reinforces the very same epistemic hierarchies Ngugi warned about decades ago.
Wanting to better understand the extent of this disconnect, I decided to run a small experiment of my own using Efik, the only local language I speak fluently. For assessment, I designed three simple benchmarks to test how well the two most popular AI systems—ChatGPT and Gemini—could understand Efik. First, could the model tested correctly translate an English sentence into Efik? Second, could it make sense of a sentence written in Efik into English? And lastly, if given a sentence out of context, could it correctly identify it was looking at Efik? The results were variable. ChatGPT, for instance, handled translations wrongly most of the time and hallucinated meanings for words. Gemini, on the other hand, fared better with translation and interpretation, but still regularly hallucinated on words it didn’t understand. On the last benchmark, both models guessed right. To stretch my experiment, I asked both ChatGPT and Gemini to tell me the tale of ‘Mutanda Eyen Namondo’, an intentional alteration of the popular Efik tale of a similar name ‘Mutanda Oyom Namondo’. The result: Gemini reported not having any knowledge about said tale and even suggested it was a fictional title. ChatGPT, however, confidently shared that the tale was a ‘beloved folktale of the Efik people’. Then it went on to tell a novel story about a girl who fell in love with a stranger and disappeared.
These results, however, come as no surprise. At present, no African language
ranks among the top 34 languages currently on the internet, and of the top languages used in AI training, none is African. Similarly, according to a paper on OpenReview, more than 60 per cent of NLP research for sub-Saharan African languages remains unfunded. And another study reveals that only two percent of Africa’s total language is present in LLMs development due to the pervasive lack of quality datasets in Africa.
Compounding these issues is the deficit in resources and infrastructure. Africa produces less than one percent of the world’s AI research and accounts for just 2.5 per cent of global AI talent, with most experts forced to leave for better opportunities abroad. Africa’s AI infrastructure, if pitted against the tightly integrated systems of the global AI giants—OpenAI, Google, Meta—who operate with deep financial reserves and continent-spanning infrastructures, appears starkly inadequate.
This disparity has led to a new kind of digital colonialism, a system where powerful AI companies feed AI systems overwhelmingly with Western-leaning datasets, often leaving out other diversities. Consequently, today’s AI is inherently biased, especially towards marginalized communities. In a study published in 2024, it was revealed that GPT-4 when prompted in African American English, tagged speakers with slurs like ‘dirty’, ‘stupid’, or ‘lazy’, revealing deep-seated prejudices dating centuries back. In image generation tests, AI that was asked for ‘Black doctors’ repeatedly placed white physicians at the centre resurrecting the old colonial ‘saviour’ trope. These are two examples of a pervasive problem; they are not glitches, but the bitter fruits of digital colonialism.
In response, African AI innovators are actively pushing for a different path: to build AI for Africa by Africans. In East and West Africa, Masakhane Project, a grassroots research collective, has brought together hundreds of African NLP researchers to create open-source models and datasets for low-resource African languages. In South Africa, Lelepa AI is working to build multilingual AI tools specifically for African users, with a focus on inclusivity, privacy, and linguistic diversity. In Nigeria particularly, its communications minister, Bosun Tijani, recently established a National Centre for Artificial Intelligence and Robotics, which, he stated, is poised to build the country’s first multilingual LLM. The African Languages Lab has used AI to amass over 400GB of speech and text data for forty low-resource languages, building the corpora that underpin new transcription, translation, and learning apps aimed at preventing further language loss. Tools like Lesan AI, trained specifically for Ethiopian languages, have outperformed larger models, further proving that with the right focus and investment Africa can fight its way towards digital independence.
Given UNESCO’s warning, for instance, that close to 500 of Africa’s 2000+ languages are endangered, it becomes clear why these initiatives are relevant. Fortunately, African leaders and tech innovators are stepping in. In April 2025, at the Global AI Summit on Africa held in Kigali, hundreds of delegates signed the ‘Africa Declaration on Artificial Intelligence’, agreeing to foundational commitments to promote the development of National AI strategies as well as to establish governance frameworks in line with the African Union’s AI strategy. To support this, present nations agreed to a pledge to invest in a $60 billion fund towards building a pan-African AI ecosystem. It is hoped that this collective effort will not only drive momentum but also ensure that Africa’s diversity is actively represented in AI development. What matters now, as Ngũgĩ, if present, would insist, is an unwavering execution.
When Ngũgĩ wrote Decolonizing the Mind decades ago, railing against the enduring grip of colonial languages on African thought, he might not have foreseen this playing out in a digital world. Though he lived to witness the AI revolution, there’s no public record on his thoughts about it. And as much as it would have been a blessing to his insights, the conversation now falls on us. So here, in his absence, the task is to ensure that Africa’s languages do not become invisible in the AI driven future. It is our duty to ensure that its languages are properly represented. To decolonize the machines, we must go beyond using our languages to actively embed them in the datasets and infrastructures shaping artificial intelligence. Africa must not merely adapt to AI; it must help define it. And that just as Ngũgĩ’ss generation fought for political and cultural sovereignty, our generation must confront the fight against digital colonialism. Because if we do not train the machines of the future in our languages and ways, our truths will be distorted and our tongues buried. Because what richer future could there be than one where AI reflects the full, vibrant spectrum of human expression, from New York to Tokyo to Ibadan, in every language under the sun?
*Prosper Ishaya emerged as the winner of the 2025 edition of The Republic Student Writing Competition, in partnership with the Open Society Foundation.
BIBLIOGRAPHY
1. Jessica Ojo et al.: Afrobench: How Good Are Large Language Models on African Languages? Published on arXiv on September 7, 2023.
3. Emily M. Bender, Timnit Gebru, Angelina McMillan-Major & Margaret Mitchell: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? Published in the ACM Conference on Fairness, Accountability, and Transparency (FAccT) on March 3, 2021.
4. The NLP Africa Community: The Current State of NLP in Sub-Saharan Africa – A Position Paper. Published on OpenReview on February 22, 2024.
5. The NLP Africa Community: The Current State of NLP in Sub-Saharan Africa – A Position Paper. Published on OpenReview on February 22, 2024.
6. Compiler News: The World Is Missing Out on Mining Something Better than Diamonds. Published on
Compiler.news on December 12, 2024.
7. Springer AI Ethics: Youth Language and Emerging Slurs: The Racial Bias of AI Classifiers on African-American English. Published in Springer AI Ethics Journal on January 2025.
8. Essence News Staff: Is AI Racist? A Researcher Tried To Generate Images Of Black Doctors. Published on
Essence.com on July 15, 2023.
9. Masakhane Project: Masakhane Is a Grassroots Organisation for NLP Research for Africans, by Africans. Published on
Masakhane.io on August 10, 2023.
10. Lerato Mdluli: Lelapa AI Launches Africa’s First AI Large Language Model, InkubaLM. Published on African World Report on October 24, 2024.
11. The Guardian Nigeria: Tijani Explains Re-Launch of Nigeria’s National Centre for AI and Robotics (NCAIR). Published by The Guardian Nigeria on September 21, 2023.
12. Smartling Blog: How the African Languages Lab Is Empowering Low-Resource Languages. Published by Smartling on April 4, 2024.
13. Andrew Deck: The AI Startup Outperforming Google Translate in Ethiopian Languages. Published on Rest of World on March 20, 2023.
14. UNESCO: Recommendation on the Safeguarding of Indigenous Languages in AI. Published by UNESCO Digital Futures Series on October 5, 2024.
15. Global AI Summit: Africa Declaration on Artificial Intelligence Signed by Delegates in Kigali. Published on
africanAIconf.org on April 12, 2025.
Comments
Sign in to comment.