datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Luhya-ASR-Data-subset-642H
Luhya ASR Data Subset 642H
Luhya speech dataset for automatic speech recognition.
tlott-digital-products
T. Lott Digital Products
Digital product files for T. Lott's online store.
Products
Audiobooks (MP3)
eBooks (PDF)
Software (ZIP)
Cover images (PNG)
Download URLs
Files can be downloaded directly:
https://huggingface.co/datasets/ziggylott/tlott-digital-products/resolve/main/{filepath}
Somali-ASR-Subset-68H
Somali ASR Subset 68H
Somali speech dataset for automatic speech recognition.
khmer-speech-dataset
Khmer ASR Cultural Dataset
727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset.Kamba-ASR-Data-Subset-484H
Kamba ASR Data Subset 484H
Kamba speech dataset for automatic speech recognition.
Gusii-ASR-Data-Subset-470H
Gusii ASR Data Subset 470H
Gusii speech dataset for automatic speech recognition.
khm-asr-cultural
Khmer ASR Cultural Dataset
134.6 hours manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8.54 seconds with the standard deviation of 3.37. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (4 females, 4 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s):… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khm-asr-cultural.digit_mask_false_positive_cv12_rawafrispeak_kinyarwanda_male_tts_datasetdigit_mask_augmented_raw
Dataset Card for "digit_mask_augmented_raw"
More Information needed
in-the-grooveCompiled from several different sets of songs:
(ITG) In the Groove
(ITG) In the Groove 2
Songs were downloaded from https://search.stepmaniaonline.net/packs/in+the+groove and are stored here for persistence.
In The Groove/ITG typically refers to DDR beatmaps done with an eye towards pad play.
Dataset info: https://paperswithcode.com/dataset/itg
Luhya-ASR-Data-subset-50hspeak-the-digit
Speak the Digit: Spoken Digit Recognition
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
2,400
Labeled training data
test
600
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
audio
Audio
id
string
label
string (test sentinel: unlabeled)… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/speak-the-digit.audios-lingala-annotatees
Annotated Lingala Dataset – Full Version
Description
This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models.
It includes:
the original audio files (viewable directly in the Hugging Face viewer)
text transcriptions
Mel spectrograms
tokenized labels
Overall statistics
Metric
Value
Total volume
5 h 0 min 18 s
Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.audios-lingala-annotatees-v2
Annotated Lingala Audio — canonical corpus
Annotated Lingala speech for open automatic speech recognition research and for
fine-tuning speech models.
This release is a full reconstruction of the corpus from its source
recordings and annotations. It supersedes
Congo-digital-service/audios-lingala-annotatees,
which is deprecated — see Relationship to the previous release below.
What this dataset contains
Each row is one annotated speech segment, carrying the audio… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v2.Luhya-ASR-Data-subset-LWLuhya-ASR-Data-subset-CAAfrivoice_Swahili-Voice_Instruct_Formatfree-spoken-digit-datasetLuhya-ASR-Data-subset-TOLuhya-ASR-Data-subsetafrispeak_kinyarwanda_female_tts_datasetDigitalUmuganda_AfriVoice_shonaLuhya-ASR-Data-subset-LAfaso-speech-dioula-digits
Faso Speech Dioula Digits
A spoken-digit classification dataset in Dioula (Jula/Bambara), built from the
Zenodo record 8320370 archive
(DOI: 10.5281/zenodo.8320370).
Each clip is one speaker saying a single digit, 1 through 4, in Dioula.
Recordings vary in speaker, accent, and recording environment.
Dataset Summary
Split
Rows
Duration
Per-class rows
train
1,532
01:34:14.9
383 / 383 / 383 / 383
validation
168
00:10:08.2
42 / 42 / 42 / 42
The split… See the full description on the dataset page: https://huggingface.co/datasets/madoss/faso-speech-dioula-digits.spoken-arabic-digits
Overview
This dataset contains spoken Arabic digits from 40 speakers from multiple Arab communities and local dialects. It is augmented using various techniques to increase the size of the dataset and improve its diversity. The recordings went through a number of pre-processors to evaluate and process the sound quality using Audacity app.
Dataset Creation
The dataset was created by collecting recordings of the digits 0-9 from 40 speakers from different Arab communities… See the full description on the dataset page: https://huggingface.co/datasets/mohnasgbr/spoken-arabic-digits.Free-Spoken-Digit-Datasetindonesia-earlymedia-digits-v1
Dataset Card for "indonesia-earlymedia-digits-v1"
More Information needed
train_valid_digit_mask_augmented_raw
Dataset Card for "train_valid_digit_mask_augmented_raw"
More Information needed
digital-love-dance
Digital love dance • Reachy Mini Moves
Community-contributed Marionette recordings captured on Reachy Mini.
Files live under data/, each move ships as a JSON trajectory plus an optional WAV.
Recorded with the Marionette web app.
Reuse
Cite this dataset as BradyM14/digital-love-dance.
Keep the reachy_mini_community_moves tag when sharing derivatives so the community can discover related sets.
