datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
distilled-catalan-youtube-speechThe Distilled Catalan YouTube Speech Corpus is the result of the automatic validation of the original
Catalan YouTube Speech Corpus created by SoftCatala and shared trhough this HF repo:
https://huggingface.co/datasets/softcatala/catalan-youtube-speech
The corpus is -distilled- because only recordings with a high certaintity to be correct were taken
and the rest were rejected.distilled-yodas-spanishDistilled YODAS Spanish is a high-quality subset of the Spanish portion of the YouTube-Oriented Dataset for Audio and Speech (YODAS). While the full YODAS corpus contains over 37,000 hours of Spanish speech across 43 million files, this dataset provides a distilled version of approximately 8,000 validated hours.nepali-oov-distilled
Nepali OOV-distilled subset (854 h)
An OOV-dense distillation of Premal-12/c9nepali-audio-dataset2 (used with the
author's permission), shipped in four variants: the original single-voice audio,
a CPU-augmented copy, and 244 h re-rendered onto 1,842 real human speakers with
Seed-VC. For Nepali ASR and TTS work.
Filter with the variant field -- see Composition below. If you came here
for speaker diversity, you want variant == "vc".
What this is
The source corpus is… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-oov-distilled.transatlantic-voice-archive_distille
Distillation brucemacd/transatlantic-voice-archive
Dataset ASR distillé via Cohere Transcribe.
Source : brucemacd/transatlantic-voice-archive
Modèle ASR : cohere-transcribe
Langue ASR : en
Exemples : 1427 (dataset source intégral)
Colonnes :
audio — clip audio (16 kHz)
transcription_base — référence brute du dataset source
transcription_cohere — hypothèse Cohere brute
langue_accent — langue / accent détecté
wer, cer — métriques item (normalisation training_v3, textes stockés… See the full description on the dataset page: https://huggingface.co/datasets/Zeldeo/transatlantic-voice-archive_distille.
