UBC-NLP/Simba-S
5249
1---2language:3 - am # Amharic4 - ar # Arabic5 - tw # Asante Twi6 - bm # Bambara7 - fr # French8 - lg # Ganda9 - ha # Hausa10 - ig # Igbo11 - rw # Kinyarwanda12 - kg # Kongo13 - ln # Lingala14 - lu # Luba-Katanga15 - mg # Malagasy16 - nso # Northern Sotho17 - ny # Nyanja18 - om # Oromo19 - pt # Portuguese20 - sn # Shona21 - so # Somali22 - st # Southern Sotho23 - sw # Swahili24 - ss # Swati25 - ti # Tigrinya26 - ts # Tsonga27 - tn # Tswana28 - ak # Twi29 - ve # Venda30 - wo # Wolof31 - xh # Xhosa32 - yo # Yoruba33 - zu # Zulu34 - tzm # Tamazight35 - sg # Sango36 - din # Dinka37 - ee # Ewe38 - fo # Fon39 - luo # Luo40 - mos # Mossi41 - umb # Umbundu42license: cc-by-4.043tags:44 - automatic-speech-recognition45 - audio46 - speech47 - african-languages48 - multilingual49 - simba50 - low-resource51 - speech-recognition52 - asr53datasets:54 - UBC-NLP/SimbaBench55metrics:56 - wer57 - cer58library_name: transformers59pipeline_tag: automatic-speech-recognition60---61<div align="center">62 63<img src="https://africa.dlnlp.ai/simba/images/VoC_simba" alt="VoC Simba Models Logo">64 65 66[](https://aclanthology.org/2025.emnlp-main.559/)67[](https://africa.dlnlp.ai/simba/)68[](https://huggingface.co/spaces/UBC-NLP/SimbaBench)69[](https://github.com/UBC-NLP/simba)70[](https://huggingface.co/collections/UBC-NLP/simba-speech-series)71[](https://huggingface.co/datasets/UBC-NLP/SimbaBench_dataset)72 73</div>74 75## *Bridging the Digital Divide for African AI*76 77**Voice of a Continent** is a comprehensive open-source ecosystem designed to bring African languages to the forefront of artificial intelligence. By providing a unified suite of benchmarking tools and state-of-the-art models, we ensure that the future of speech technology is inclusive, representative, and accessible to over a billion people.78 79## Best-in-Class Multilingual Models80 81Introduced in our EMNLP 2025 paper *[Voice of a Continent](https://aclanthology.org/2025.emnlp-main.559/)*, the **Simba Series** represents the current state-of-the-art for African speech AI.82 83- **Unified Suite:** Models optimized for African languages.84- **Superior Accuracy:** Outperforms generic multilingual models by leveraging SimbaBench's high-quality, domain-diverse datasets.85- **Multitask Capability:** Designed for high performance in ASR (Automatic Speech Recognition) and TTS (Text-to-Speech).86- **Inclusion-First:** Specifically built to mitigate the "digital divide" by empowering speakers of underrepresented languages.87 88The **Simba** family consists of state-of-the-art models fine-tuned using SimbaBench. These models achieve superior performance by leveraging dataset quality, domain diversity, and language family relationships.89 90### ๐ฃ๏ธโ๏ธ Simba-ASR91> **The New Standard for African Speech-to-Text**92 93**๐ฏ Task** `Automatic Speech Recognition` โ Powering high-accuracy transcription across the continent.94 95**๐ Language Coverage (43 African languages)**96> **Amharic** (`amh`), **Arabic** (`ara`), **Asante Twi** (`asanti`), **Bambara** (`bam`), **Baoulรฉ** (`bau`), **Bemba** (`bem`), **Ewe** (`ewe`), **Fanti** (`fat`), **Fon** (`fon`), **French** (`fra`), **Ganda** (`lug`), **Hausa** (`hau`), **Igbo** (`ibo`), **Kabiye** (`kab`), **Kinyarwanda** (`kin`), **Kongo** (`kon`), **Lingala** (`lin`), **Luba-Katanga** (`lub`), **Luo** (`luo`), **Malagasy** (`mlg`), **Mossi** (`mos`), **Northern Sotho** (`nso`), **Nyanja** (`nya`), **Oromo** (`orm`), **Portuguese** (`por`), **Shona** (`sna`), **Somali** (`som`), **Southern Sotho** (`sot`), **Swahili** (`swa`), **Swati** (`ssw`), **Tigrinya** (`tir`), **Tsonga** (`tso`), **Tswana** (`tsn`), **Twi** (`twi`), **Umbundu** (`umb`), **Venda** (`ven`), **Wolof** (`wol`), **Xhosa** (`xho`), **Yoruba** (`yor`), **Zulu** (`zul`), **Tamazight** (`tzm`), **Sango** (`sag`), **Dinka** (`din`).97 98**๐๏ธ Base Architectures**99 100 - **Simba-S** (SeamlessM4T-v2-MT) โ *Top Performer*101 - **Simba-W** (Whisper-v3-large)102 - **Simba-X** (Wav2Vec2-XLS-R-2b)103 - **Simba-M** (MMS-1b-all)104 - **Simba-H** (AfriHuBERT)105 106๐ Explore the Frontier107 108| **ASR Models** | **Architecture** | **#Parameters** | **๐ค Hugging Face Model Card** | **Status** |109|---------|:------------------:| :------------------:| :------------------:|:------------------:| 110| ๐ฅ**Simba-S**๐ฅ| SeamlessM4T-v2 | 2.3B | ๐ค [https://huggingface.co/UBC-NLP/Simba-S](https://huggingface.co/UBC-NLP/Simba-S) | โ
Released |111| ๐ฅ**Simba-W**๐ฅ| Whisper | 1.5B | ๐ค [https://huggingface.co/UBC-NLP/Simba-W](https://huggingface.co/UBC-NLP/Simba-W) | โ
Released | 112| ๐ฅ**Simba-X**๐ฅ| Wav2Vec2 | 1B | ๐ค [https://huggingface.co/UBC-NLP/Simba-X](https://huggingface.co/UBC-NLP/Simba-X) | โ
Released | 113| ๐ฅ**Simba-M**๐ฅ| MMS | 1B | ๐ค [https://huggingface.co/UBC-NLP/Simba-M](https://huggingface.co/UBC-NLP/Simba-M) | โ
Released | 114| ๐ฅ**Simba-H**๐ฅ| HuBERT | 94M | ๐ค [https://huggingface.co/UBC-NLP/Simba-H](https://huggingface.co/UBC-NLP/Simba-H) | โ
Released | 115 116* **Simba-S** emerged as the best-performing ASR model overall.117 118 119**๐งฉ Usage Example**120 121You can easily run inference using the Hugging Face `transformers` library.122 123```python124from transformers import pipeline125 126# Load Simba-S for ASR127asr_pipeline = pipeline(128 "automatic-speech-recognition",129 model="UBC-NLP/Simba-S" #Simba mdoels `UBC-NLP/Simba-S`, `UBC-NLP/Simba-W`, `UBC-NLP/Simba-X`, `UBC-NLP/Simba-H`, `UBC-NLP/Simba-M`130)131 132##### Load the multilingual African adapter (Only for `UBC-NLP/Simba-M`)133asr_pipeline.model.load_adapter("multilingual_african") # Only for `UBC-NLP/Simba-M`134###########################135 136# Transcribe audio from file137result = asr_pipeline("https://africa.dlnlp.ai/simba/audio/afr_Lwazi_afr_test_idx3889.wav")138print(result["text"])139 140 141# Transcribe audio from audio array142result = asr_pipeline({143 "array": audio_array,144 "sampling_rate": 16_000145})146print(result["text"])147 148```149 150#### Example Outputs151 152Using the same audio file with different Simba models:153 154```python155# Simba-S156{'text': 'watter verontwaardiging sou daar, in ons binneste gewees het.'}157```158 159```python160# Simba-W161{'text': 'watter veronwaardigingsel daar, in ons binneste gewees het.'}162```163 164```python165# Simba-X166{'text': 'fator fr on ar taamsodr is'}167```168 169```python170# Simba-M171{'text': 'watter veronwaardiging sodaar in ons binniste gewees het'}172```173 174```python175# Simba-H176{'text': 'watter vironwaardiging so daar in ons binneste geweeshet'}177```178 179Get started with Simba models in minutes using our interactive Colab notebook: [](https://github.com/UBC-NLP/simba/blob/main/simba_models.ipynb)180 181 182## Citation183 184If you use the Simba models or SimbaBench benchmark for your scientific publication, or if you find the resources in this website useful, please cite our paper.185 186```bibtex187 188@inproceedings{elmadany-etal-2025-voice,189 title = "Voice of a Continent: Mapping {A}frica{'}s Speech Technology Frontier",190 author = "Elmadany, AbdelRahim A. and191 Kwon, Sang Yun and192 Toyin, Hawau Olamide and193 Alcoba Inciarte, Alcides and194 Aldarmaki, Hanan and195 Abdul-Mageed, Muhammad",196 editor = "Christodoulopoulos, Christos and197 Chakraborty, Tanmoy and198 Rose, Carolyn and199 Peng, Violet",200 booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing",201 month = nov,202 year = "2025",203 address = "Suzhou, China",204 publisher = "Association for Computational Linguistics",205 url = "https://aclanthology.org/2025.emnlp-main.559/",206 doi = "10.18653/v1/2025.emnlp-main.559",207 pages = "11039--11061",208 ISBN = "979-8-89176-332-6",209}210 211```212 213 