datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bambara-Speech-Translation-Data
AfVoices-Translated (Bambara-English)
This is a Bambara speech translation dataset, which is built on the African Next Voices (AfVoices) Bambara ASR corpus. It provides English translations for the human-corrected subset of the original collection, creating a parallel corpus for Bambara-English machine translation and speech-to-text tasks.
Methodology
We machine-translated the human-validated transcriptions from AfVoices using the Oolel-translator repository.
Inference… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Bambara-Speech-Translation-Data.AST-Music-Data-45KAST-Music-Data-82KWolof-ASR-DataA curated Wolof ASR dataset from various sources:
Split
Fleurs
Alfa
CV
Kallama
UB
Total
Train
8.72
16.13
34.97
33.60
4.52
97.94
Test
1.75
2.84
6.21
5.91
1.12
17.83
This dataset was used to finetune Wolof-HuBERT-Base for ASR.
ukr-dialects-audio-dataset
Ukrainian Dialects Audio Dataset
Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits.
Dataset Description
This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets:
NaUKMA-Audio-Dataset
Ivanna-Stefiuk-Audio-Dataset
Larysa-Irodenko-Audio-Dataset
Hutsulendia-Audio-Dataset
Dido-Yvanchyk-Audio-Dataset-v2
Dataset Structure
train: 27,675 samples
validation: 3… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukr-dialects-audio-dataset.dataset-from-restoreAST-Music-Data-1Kthaha-research-data2-v2
Nepali Speech Dataset (YouTube-sourced)
441 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 441 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2-v2.thaha-research-data1
Nepali Speech Dataset (YouTube-sourced)
59 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 59 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data1.thaha-research-data2
Nepali Speech Dataset (YouTube-sourced)
63 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 63 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2.thaha-research-data2-v2-v2
Nepali Speech Dataset (YouTube-sourced)
76 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 76 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2-v2-v2.AST-Music-Data-291K
Speech-Music Merged Dataset
Source Datasets
Dataset
Samples
Speech
Music
AIGenLab/high-sound-and-low-music
91,054
0
91,054
AIGenLab/Speech_Dataset
100,000
100,000
0
AIGenLab/Music-Dataset
100,000
0
100,000
TOTAL
291,054
100,000
191,054
Dataset Info
Total Samples: 291,054
Labels: speech, music
Audio: 16kHz, Mono, WAV
Usage
from datasets import load_dataset
dataset = load_dataset("AIGenLab/speech-music-merge", split="train")… See the full description on the dataset page: https://huggingface.co/datasets/Vyvo-Research/AST-Music-Data-291K.ASR_datasetecho-synthetic-data-10kresearch_dataset_v1
