CoolFace
Datasetpublic

ciatech-frica/central-kanuri-speech-dataset

anguage: kau license: cc-by-nc-4.0 pretty_name: Central Kanuri Speech Dataset task_categories: automatic-speech-recognition tags: kanuri central-kanuri speech audio asr low-resource-language african-languages speech-recognition conversational-ai Central Kanuri Speech Dataset Overview The Central Kanuri Speech Dataset is a community-contributed speech corpus developed by CIATECH Africa in collaboration with CLEAR Global through the TWB Voice initiative. The dataset was developed to increase the… See the full description on the dataset page: https://huggingface.co/datasets/ciatech-frica/central-kanuri-speech-dataset.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes43downloads
Dataset Card

anguage:

kau license: cc-by-nc-4.0 prettyname: Central Kanuri Speech Dataset taskcategories: automatic-speech-recognition tags: kanuri central-kanuri speech audio asr low-resource-language african-languages speech-recognition conversational-ai Central Kanuri Speech Dataset Overview

The Central Kanuri Speech Dataset is a community-contributed speech corpus developed by CIATECH Africa in collaboration with CLEAR Global through the TWB Voice initiative.

The dataset was developed to increase the availability of high-quality speech resources for Central Kanuri (kau), a low-resource language spoken across communities in the Lake Chad Basin.

It is intended to support research and development in speech and language technologies, including Automatic Speech Recognition (ASR), speech-to-text, language identification, speech translation, conversational AI, and other low-resource language technologies.

Dataset Composition

The corpus contains two types of speech data:

Read Speech

Participants recorded Central Kanuri speech by reading predefined prompts.

The recordings went through a recording and recording-quality-check workflow.

Freeform Speech

Participants provided spontaneous/freeform speech in Central Kanuri.

The freeform recordings followed a workflow that included recording, recording quality checks, transcription, transcription checks, and, where applicable, transcription correction.

The Read and Freeform components are maintained separately because they were produced through different data-collection and quality-assurance workflows.

Language Language: Central Kanuri ISO 639-3: kau Language family: Saharan Collection context: Lake Chad Basin Dataset Statistics

The dataset currently contains approximately 3,576 audio records and has a repository size of approximately 5.56 GB.

The exact statistics may change as the dataset is validated and additional releases are published.

Data Collection

Data collection was conducted through the TWB Voice platform operated by CLEAR Global, with CIATECH Africa supporting implementation and community engagement.

The collection included both predefined reading prompts and spontaneous/freeform speech.

The project included contributor onboarding, recording guidance, quality review, transcription where applicable, and linguistic validation.

Data Processing and Quality Assurance

The data followed the TWB Voice workflow.

Read Speech Recording ↓ Recording Check Freeform Speech Recording ↓ Recording Check ↓ Transcription ↓ Transcription Check ↓ Optional Transcription Correction

The published data has been pseudonymised and processed to reduce the inclusion of personally identifying information.

Intended Uses

The dataset is intended to support:

Automatic Speech Recognition (ASR) Speech-to-text systems Central Kanuri language technology Language identification Speech translation Conversational AI Low-resource speech and NLP research African-language AI development Humanitarian and development-oriented language technologies Research into multilingual and low-resource speech systems Ethical Considerations and Privacy

This dataset contains recordings of human voices. Although contributor identifiers and personally identifying information have been pseudonymised or removed from the published dataset, voice recordings should be treated as potentially sensitive data.

Users should:

respect contributor privacy; not attempt to identify individual speakers; not use the dataset for speaker identification or biometric profiling; not attempt to re-identify contributors; not combine the dataset with other information for the purpose of identifying individuals; and comply with the applicable license and data-use conditions. Limitations

The dataset represents a particular collection population, geographic context, recording environment, and linguistic setting. It should not be assumed to represent all Kanuri speakers or all varieties of Kanuri.

Potential sources of variation and bias include:

contributor demographics; geographic distribution; recording devices and environments; speaking styles; prompt design; linguistic variation; differences between read and spontaneous speech; and dataset size.

Researchers should evaluate these factors when using the dataset to train or evaluate speech and language models.

License

This dataset is released under the Creative Commons Attribution–NonCommercial 4.0 International (CC BY-NC 4.0) license.

Users must comply with the terms of the license and provide appropriate attribution.

Attribution

If you use this dataset in research, publications, models, demonstrations, or other work, please acknowledge:

CIATECH Africa CLEAR Global TWB Voice the participating Kanuri-speaking contributors and language experts the validators and transcribers who supported the dataset development Suggested citation @dataset{grema2026centralkanurispeech, title={Central Kanuri Speech Dataset}, author={Grema, Umar and CIATECH Africa and CLEAR Global}, year={2026}, publisher={Hugging Face}, license={CC BY-NC 4.0} } Acknowledgements

CIATECH Africa and CLEAR Global acknowledge the Kanuri-speaking community members who contributed recordings, as well as the language experts, validators, transcribers, coordinators, and other contributors who supported the development and quality assurance of this corpus.

We also acknowledge the TWB Voice team for providing the data-collection and workflow infrastructure supporting this project.

Versioning

Future releases may include additional validated recordings, improved transcriptions, metadata corrections, or expanded linguistic coverage.

Users should reference the specific dataset version used in their research or model development.

Contact

CIATECH Africa

Website: https://ciatech.ng

Project Lead: Umar Grema