CoolFace
Datasetpublic

scubavoice/ibibio-efik-speech-corpus-sample

Scuba Voice Dataset: Ibibio & Efik Sample (1 hour) Scuba is a voice data infrastructure company collecting, labeling, and licensing speech corpora for low-resource languages. This repository contains a 1-hour sample of labeled Ibibio and Efik speech data collected from native speakers in Akwa Ibom, Nigeria. This sample is provided for evaluation purposes only. For access to the full dataset (~200 hours across Ibibio and Efik), contact us at info@scubavoice.com.… See the full description on the dataset page: https://huggingface.co/datasets/scubavoice/ibibio-efik-speech-corpus-sample.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
1likes25downloads
Dataset Card

Scuba Voice Dataset: Ibibio & Efik Sample (1 hour)

Scuba is a voice data infrastructure company collecting, labeling, and licensing speech corpora for low-resource languages. This repository contains a 1-hour sample of labeled Ibibio and Efik speech data collected from native speakers in Akwa Ibom, Nigeria.

This sample is provided for evaluation purposes only. For access to the full dataset (~200 hours across Ibibio and Efik), contact us at info@scubavoice.com.


Languages

FieldIbibioEfik
ISO 639-3ibbefi
RegionAkwa Ibom State, NigeriaAkwa Ibom & Cross River State, Nigeria
Language familyNiger-Congo, Cross RiverNiger-Congo, Cross River
Speakers (estimated)~7 million~3 million

Both Ibibio and Efik are tonal languages. This dataset includes prompt-type labels that reflect the emotional register of each utterance, which is relevant for tone-sensitive Speech-to-Intent, ASR and TTS modeling.


Dataset Structure

scuba-ibibio-efik-sample/
├── audio/               # .webm audio files
├── metadata.csv         # Per-clip metadata
└── README.md

Metadata fields

ColumnDescription
file_namePath to audio file
recording_idUnique clip identifier
languageLanguage name
iso_639_3ISO 639-3 language code
english_translationEnglish gloss of the prompt
prompt_typeEmotional register (neutral, serious, worried, etc.)
speaker_idAnonymized speaker identifier
genderSpeaker gender (self-reported)
age_rangeSpeaker age bracket
duration_secondsClip duration
licenseLicense type
consent_statusConsent collection method
consent_versionConsent form version

Sample Statistics

MetricValue
Total duration~1 hour
LanguagesIbibio (ibb), Efik (efi)
Full dataset100+ hours, Ibibio + Efik

License

This sample is released under the [Scuba Sample Voice Data License](https://huggingface.co/datasets/scubavoice/ibibio-efik-speech-corpus-sample/resolve/main/LICENSE).

  • Free for evaluation and non-commercial research
  • Commercial use requires a separate license agreement
  • Redistribution is not permitted

To license the full dataset: info@scubavoice.com


Full Dataset Access

This sample represents <1% of Scuba's current corpus. The full dataset includes:

  • ~200 hours of labeled speech
  • Ibibio and Efik (ISO: ibb, efi)
  • 500+ speakers across gender and age groups
  • Dialect coverage across South South Nigeria
  • Conversational english prompt text with emotional register labels (neutral, serious, worried, etc.)

For licensing, research partnerships, or custom collection: info@scubavoice.com


Citation

If you use this data in published research, please cite:

@dataset{scuba2026ibibio_efik,
  title     = {Scuba Ibibio-Efik Voice Dataset},
  author    = {Scuba},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/datasets/scuba-ai/ibibio-efik-sample}
}

About Scuba

Scuba collects, labels, and licenses speech corpora for languages that AI has left behind. We work directly inside speaker communities to build the data infrastructure that makes voice AI possible for everyone.

Website: scubavoice.com Contact: info@scubavoice.com