When Vambo AI, a Johannesburg-based artificial intelligence (AI) model builder, set out to build MORENA, a 1.5-billion-parameter model covering 12 African languages, it had access to a supercomputer through a United Nations Development Programme (UNDP)-backed programme. It had the graphics processing units (GPUs), specialised chips that power AI computing, but struggled to find enough real African-language text.
“Finding that much real data was close to impossible, so a lot of it was synthetic,” Isheanesu Misi, Vambo’s co-founder and chief technology officer (CTO), told TechCabal in an interview on Monday.
The shortage of African-language text data threatens to limit the development of AI tools that reflect the continent’s linguistic diversity. Without enough high-quality data, developers risk building AI products that perform poorly in local languages, leaving millions of potential users underserved.
MORENA was trained on 251.7 billion tokens, including 65.6 billion African-language tokens, according to its model documentation. The model covers languages such as ChiShona, Kiswahili, Hausa, Yorùbá, Igbo, isiZulu, isiXhosa, Kinyarwanda, Setswana, Afrikaans, isiNdebele and Nigerian Pidgin, alongside English, French and code.
Vambo says the model was trained from scratch rather than adapting an existing foundation model because models developed on general-purpose bases can inherit weaknesses in how they handle African languages.
The result is a model that the company says performs strongly on African-language text modelling. MORENA recorded 1.408 bits per byte across the 12 languages, the lowest score among 26 models tested, including the 8-billion-parameter Lugha-Llama-8B and larger general-purpose models.
Its tokeniser, the system that breaks text into pieces for an AI model to process, also represents African-language text using 1.39 times fewer tokens than Gemma 3 and 1.53 times fewer than Llama 3.2, according to Vambo’s model card.
Misi said synthetic data helped fill the gap left by scarce real-world text, but it introduced imperfections. “You would find that some of the grammar of the outputs is close but imperfect,” he said. Synthetic data also helped steer the model towards specific capabilities, he added, including tool calling.
The shortage is particularly striking because some of the languages involved have millions of speakers. Misi said languages such as isiNdebele and Nigerian Pidgin can have “next to no resources available” for AI development.
“That means that the AI applications are extremely important and necessary, but they can’t be fully built by local innovators because they don’t have everything they need,” he said.
Compute presented another barrier, but Vambo overcame it through outside support. The company did not pay for the GPUs used to train MORENA. CINECA, an Italian inter-university consortium specialising in high-performance computing and research, provided access to its Leonardo supercomputer through the UNDP AI Hub for Sustainable Development. Some training runs used 512 GPUs simultaneously.
“Without access to that kind of compute power, this project would have been dead on arrival,” Misi said.
The 22,000 A100 GPU-hours used to train MORENA were estimated to be worth roughly $25,000- $ 40,000, but Vambo’s cash cost was zero because the compute was provided by CINECA through the UNDP AI Hub for Sustainable Development.
According to Misi, the project demonstrates what African teams can build when high-performance infrastructure is made available, but also raises questions about how repeatable such projects are without subsidised access.
For now, Vambo is positioning MORENA as infrastructure rather than a finished consumer product. Its weights are available under an Apache 2.0 licence, while the training corpus has not yet been released. Smaller 0.5-billion and 0.2-billion versions are designed for lower-resource environments.
Misi says developers could fine-tune MORENA for applications such as fintech or education without bearing the cost of training a foundation model from scratch. “If everyone is to try and build a model from scratch, it’s prohibitively expensive, complicated, time-consuming,” Misi said. “But fine-tuning could cost a hundred dollars or even less at times.”
MORENA is part of a broader push to build AI systems that work better with African languages. The continent is home to more than 2,000 languages, yet many have limited digitised text and speech data for training AI models.
TechCabal reported in February that researchers and startups are using approaches such as voice-data collection and synthetic data to fill those gaps, while governments and Big Tech are also investing in developing African-language AI.
Nigeria’s N-ATLAS, a multilingual AI model, for example, is being developed around local languages, while Google’s WAXAL project has created an open-source speech dataset covering 21 sub-Saharan African languages.
Those efforts point to a broader challenge for African AI. Compute is becoming more accessible through initiatives such as Vambo’s CINECA-backed programme, while Nigeria is also building its own GPU infrastructure and semiconductor capabilities through efforts involving UduTech, a GPU cloud and infrastructure platform, and Chipmango, a semiconductor tech company that designs chips and trains engineers. Earlier this month, UduTech partnered with South Korean AI infrastructure company BARO AI to expand the supply of high-performance GPUs to African governments, businesses and research institutions.
For developers, however, access to computing power is only one part of the equation. Language data must also be collected, cleaned, documented, and made available at scale, making the data layer just as important to Africa’s AI ambitions as the GPUs used to train the models.
True scale demands moving beyond surface-level integrations to robust execution. We’ve filtered the noise out of Moonshot 2026, optimising the conference strictly for high-calibre connections between startup founders, global financial operators, enterprise leaders and individuals rewiring Africa’s technical frameworks. Get 20% off Early Bird tickets for a limited time.
