datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Luhya-ASR-Data-subset-642H
Luhya ASR Data Subset 642H
Luhya speech dataset for automatic speech recognition.
Somali-ASR-Subset-68H
Somali ASR Subset 68H
Somali speech dataset for automatic speech recognition.
khmer-speech-dataset
Khmer ASR Cultural Dataset
727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset.Kamba-ASR-Data-Subset-484H
Kamba ASR Data Subset 484H
Kamba speech dataset for automatic speech recognition.
Gusii-ASR-Data-Subset-470H
Gusii ASR Data Subset 470H
Gusii speech dataset for automatic speech recognition.
khm-asr-cultural
Khmer ASR Cultural Dataset
134.6 hours manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8.54 seconds with the standard deviation of 3.37. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (4 females, 4 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s):… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khm-asr-cultural.audios-lingala-annotatees
Annotated Lingala Dataset – Full Version
Description
This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models.
It includes:
the original audio files (viewable directly in the Hugging Face viewer)
text transcriptions
Mel spectrograms
tokenized labels
Overall statistics
Metric
Value
Total volume
5 h 0 min 18 s
Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.Afrivoice_V2
Dataset Card for the image text and voice dataset
Dataset Description
This is an image prompt ASR dataset 4 African languages; Kirundi, Ndau, Ndebele, and Oshiwambo. Each language was collected on at most 2 domains.
Language
Domain
Total hours
Total transcribed hours
Total clips
Size (GB)
Kirundi
Agriculture
125.42
125.42
24,926
7.2327
Education
383.66
383.66
72,994
24.6425
Sub Total
509.08
509.08
97,920
31.8752
Ndau
Education
90.60
90.60
17… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/Afrivoice_V2.Afrivoice_Swahili
Dataset Card for the image text and voice dataset
Dataset Description
This is an image prompt ASR dataset for Swahili. The dataset was collected on 5 domains: Agriculture, Education, Finance, Government and Health.
Domain
Total Hours
Transcribed Hours
Total Clips
Dataset Size (GB)
Agriculture
766.21
744.20
132,848
45.79
Education
565.74
549.52
98,246
55.25
Financial
616.05
605.54
106,500
60.83
Government
591.40
579.41
102,961
49.99
Health… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/Afrivoice_Swahili.Afrivoice_Kinyarwanda
Dataset Card for the image text and voice dataset
Dataset Description
Each datapoint in this dataset consists of a JPEG image, a corresponding audio Webm file describing the image, and when available, the transcription of the audio file.
Domain
Total Hours
Transcribed Hours
Number of Clips
Dataset Size (GB)
Agriculture
467.13
465.40
86,305
30.13
Health
994.32
992.87
179,219
58.53
Finance
564.21
563.11
103,159
38.55
Government
676.10
674.11
122,265
49.22… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/Afrivoice_Kinyarwanda.audios-lingala-annotatees-v2
Annotated Lingala Audio — canonical corpus
Annotated Lingala speech for open automatic speech recognition research and for
fine-tuning speech models.
This release is a full reconstruction of the corpus from its source
recordings and annotations. It supersedes
Congo-digital-service/audios-lingala-annotatees,
which is deprecated — see Relationship to the previous release below.
What this dataset contains
Each row is one annotated speech segment, carrying the audio… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v2.Afrivoice_Ethiopia
Afrivoice Ethiopia
Dataset Summary
Language
Category
Total number of hours
Total number of transcribed hours
Total number of clips
Total of Size of the dataset in GB
Amharic
Unscripted
351.17
81.67
71119
20.7
Scripted
113.55
113.55
22412
7.9
Expert
153.44
24.28
31144
8.1
Afaan OromoUnscripted
338.36
80.44
65555
21.8
Scripted
111.93
111.93
23051
9.7
Expert
153.7
29.88
29638
8.2
Sidama
Unscripted
348.67
74.94
72801
23.5
Scripted
114.93
114.93… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/Afrivoice_Ethiopia.Afrivoice_Swahili_0.0
Dataset summary
Domain
Total number of hours
Total number of transcribed hours
Total number of clips
Total Size of the dataset in GB
Agriculture
35.19
35.19
5,852
2.6
Health
62.11
62.11
10,315
6.4
Finance
118.75
118.75
19,629
16.6
Government
106.41
106.41
17,530
12.1
Education
91.96
91.96
15,204
6.2
Total
414.42
414.42
68,530
43.9
How to use
The datasets library allows you to load and pre-process your dataset in pure Python, at scale. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/Afrivoice_Swahili_0.0.Afrivoice_Kinyarwanda_old_version
Dataset summary
[need more information]
Supported tasks
[need more information]
How to use
[need more information]
Dataset structure
Data fields
[need more information]
Data splits
[need more information]
Data preprocessing
[need more information]
Licensing Information
All datasets are licensed under the Creative Commons license (CC-BY-4).
