
The market dominance of such LLMs puts developing countries with minority languages at a severe disadvantage.
Ever since OpenAI introduced the ChatGPT chatbot in November 2022, tech industry leaders and policymakers have been touting the potential of AI to improve society through revolutions in healthcare and education, as well as by boosting productivity and income. But these benefits are distributed unevenly.
Many developing countries cannot access AI systems due to their linguistic bias. Most LLMs are trained on approximately 100 of the more than 7,000 languages spoken in the world today.
English and other “high-resource” languages of the “Big Five” (Chinese, Spanish, French, and German) account for nearly all of the data used. However, English is the native language of only 400 million people, and only 25% of the world’s population speaks one of the “Big Five” languages as their first language.
Such an imbalance could lead to cultural and economic isolation in the short term, and in the long term, it could slow economic growth and reduce linguistic diversity.
The bitter irony is that the rapid adoption of AI in wealthy countries, supported by the large-scale construction of data centers with powerful computing infrastructure, makes it difficult for users in developing countries to access the internet and the benefits of digital tools. The growing demand for chips driven by AI has driven up the price of smartphones, making them unaffordable for millions of people.
A lack of diversity will slow the adoption of AI systems
Large language models must better account for local languages and cultures. Otherwise, societies risk losing their unique lexicon, which has long been shaped by historical events, mass migration, and political heritage.
Moreover, the current lack of linguistic diversity in AI systems will slow their adoption. Digital financial services in Peru must be provided in simple and understandable language for consumers to appreciate them. And in Vietnam, effective online medical treatment requires a level of clarity that only a model in the native language can provide.
Mobile operators are uniquely positioned to create local language models, as they have entire teams of programmers and the capacity to process data. In addition, they hold the necessary licenses and enjoy the trust of the authorities as partners.
This is important because one of the main barriers to creating LLMs for small languages is the lack of training data. The solution could lie in the vast amounts of data held by governments, which is typically confidential.
For example, Kyivstar, Ukraine’s largest mobile operator, has formed a partnership with the Ministry of Digital Transformation to develop the country’s first national LLM, adapted to the Ukrainian language and other languages spoken in the country, including Russian, Bulgarian, and Crimean Tatar. The task was complicated by the fact that a significant portion of Ukraine’s literature was written during the period of Russian rule and was therefore deemed culturally unacceptable. As a result, the selection of training texts was overseen by four advisory committees empowered to determine the model’s technical, legal, cultural, historical, and linguistic aspects.
Developing Countries with AI Models
The Global System for Mobile Communications Association (GSMA), of which I am the CEO, is pursuing a similar strategy in many sub-Saharan African countries, including Nigeria, Madagascar, Togo, and the Democratic Republic of the Congo. This is the“African AI Language Models”initiative, which brings together not only mobile operators from the “Group of Six” (Airtel, Axian Telecom, Ethio Telecom, MTN, Orange, and Vodacom) but also other telecommunications companies, public and private organizations, researchers, and civil society groups to create scalable, inclusive large language models (LLMs) for African languages.
To foster the growth of AI startups that deliver positive socioeconomic outcomes for underserved populations, the GSMA launched the EmergingTech program. It is funded by the UK Department for Foreign, Commonwealth, and Development Affairs.
Among the companies in the first cohort are ToumAI, which is developing voice interfaces in local Moroccan dialects to help low-literacy rural users access digital and financial services using simple voice commands, and Wiseyak, which is developing a multilingual, mixed-code conversational AI solution to reduce the digital divide in Nepal.
These initiatives are crucial as the world transitions from AI that responds to requests to autonomous, agent-based AI. In digital services—including banking, healthcare, and education—artificial intelligence is increasingly being used to assist users and will become a key factor in sustaining livelihoods and incomes. It is also worth noting that open-source models are beneficial for companies.
Investments in local AI models create a virtuous cycle in developing countries, as they build a pool of technical talent—which is crucial for digital independence.
Amid growing concerns about data sovereignty and data processing, developing countries must improve operational oversight and governance when implementing agent-based AI. This means working with developers and mobile network operators to create models in local languages that meet the needs of communities in developing countries.

Vivek Badrinath,
CEO of the GSMA.
© Project Syndicate, 2026.
www.project-syndicate.org



















