ciatech-frica/central-kanuri-speech-dataset
anguage: kau license: cc-by-nc-4.0 pretty_name: Central Kanuri Speech Dataset task_categories: automatic-speech-recognition tags: kanuri central-kanuri speech audio asr low-resource-language african-languages speech-recognition conversational-ai Central Kanuri Speech Dataset Overview The Central Kanuri Speech Dataset is a community-contributed speech corpus developed by CIATECH Africa in collaboration with CLEAR Global through the TWB Voice initiative. The dataset was developed to increase the… See the full description on the dataset page: https://huggingface.co/datasets/ciatech-frica/central-kanuri-speech-dataset.
anguage:
kau license: cc-by-nc-4.0 prettyname: Central Kanuri Speech Dataset taskcategories: automatic-speech-recognition tags: kanuri central-kanuri speech audio asr low-resource-language african-languages speech-recognition conversational-ai Central Kanuri Speech Dataset Overview
The Central Kanuri Speech Dataset is a community-contributed speech corpus developed by CIATECH Africa in collaboration with CLEAR Global through the TWB Voice initiative.
The dataset was developed to increase the availability of high-quality speech resources for Central Kanuri (kau), a low-resource language spoken across communities in the Lake Chad Basin.
It is intended to support research and development in speech and language technologies, including Automatic Speech Recognition (ASR), speech-to-text, language identification, speech translation, conversational AI, and other low-resource language technologies.
Dataset Composition
The corpus contains two types of speech data:
Read Speech
Participants recorded Central Kanuri speech by reading predefined prompts.
The recordings went through a recording and recording-quality-check workflow.
Freeform Speech
Participants provided spontaneous/freeform speech in Central Kanuri.
The freeform recordings followed a workflow that included recording, recording quality checks, transcription, transcription checks, and, where applicable, transcription correction.
The Read and Freeform components are maintained separately because they were produced through different data-collection and quality-assurance workflows.
Language Language: Central Kanuri ISO 639-3: kau Language family: Saharan Collection context: Lake Chad Basin Dataset Statistics
The dataset currently contains approximately 3,576 audio records and has a repository size of approximately 5.56 GB.
The exact statistics may change as the dataset is validated and additional releases are published.
Data Collection
Data collection was conducted through the TWB Voice platform operated by CLEAR Global, with CIATECH Africa supporting implementation and community engagement.
The collection included both predefined reading prompts and spontaneous/freeform speech.
The project included contributor onboarding, recording guidance, quality review, transcription where applicable, and linguistic validation.
Data Processing and Quality Assurance
The data followed the TWB Voice workflow.
Read Speech Recording ↓ Recording Check Freeform Speech Recording ↓ Recording Check ↓ Transcription ↓ Transcription Check ↓ Optional Transcription Correction
The published data has been pseudonymised and processed to reduce the inclusion of personally identifying information.
Intended Uses
The dataset is intended to support:
Automatic Speech Recognition (ASR) Speech-to-text systems Central Kanuri language technology Language identification Speech translation Conversational AI Low-resource speech and NLP research African-language AI development Humanitarian and development-oriented language technologies Research into multilingual and low-resource speech systems Ethical Considerations and Privacy
This dataset contains recordings of human voices. Although contributor identifiers and personally identifying information have been pseudonymised or removed from the published dataset, voice recordings should be treated as potentially sensitive data.
Users should:
respect contributor privacy; not attempt to identify individual speakers; not use the dataset for speaker identification or biometric profiling; not attempt to re-identify contributors; not combine the dataset with other information for the purpose of identifying individuals; and comply with the applicable license and data-use conditions. Limitations
The dataset represents a particular collection population, geographic context, recording environment, and linguistic setting. It should not be assumed to represent all Kanuri speakers or all varieties of Kanuri.
Potential sources of variation and bias include:
contributor demographics; geographic distribution; recording devices and environments; speaking styles; prompt design; linguistic variation; differences between read and spontaneous speech; and dataset size.
Researchers should evaluate these factors when using the dataset to train or evaluate speech and language models.
License
This dataset is released under the Creative Commons Attribution–NonCommercial 4.0 International (CC BY-NC 4.0) license.
Users must comply with the terms of the license and provide appropriate attribution.
Attribution
If you use this dataset in research, publications, models, demonstrations, or other work, please acknowledge:
CIATECH Africa CLEAR Global TWB Voice the participating Kanuri-speaking contributors and language experts the validators and transcribers who supported the dataset development Suggested citation @dataset{grema2026centralkanurispeech, title={Central Kanuri Speech Dataset}, author={Grema, Umar and CIATECH Africa and CLEAR Global}, year={2026}, publisher={Hugging Face}, license={CC BY-NC 4.0} } Acknowledgements
CIATECH Africa and CLEAR Global acknowledge the Kanuri-speaking community members who contributed recordings, as well as the language experts, validators, transcribers, coordinators, and other contributors who supported the development and quality assurance of this corpus.
We also acknowledge the TWB Voice team for providing the data-collection and workflow infrastructure supporting this project.
Versioning
Future releases may include additional validated recordings, improved transcriptions, metadata corrections, or expanded linguistic coverage.
Users should reference the specific dataset version used in their research or model development.
Contact
CIATECH Africa
Website: https://ciatech.ng
Project Lead: Umar Grema
