datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Earnings22-Cleaned-AA
Earnings22-Cleaned-AA
Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article
Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.dahih-tts2-demucs-cleanedfarsi-asr-unified-cleaned
🎧 Farsi ASR Unified Dataset (Parquet Sharded Edition)
Overview
The Farsi ASR Unified Dataset is a large-scale, high-quality, and fully standardized collection of Persian (Farsi) speech-to-text data — designed specifically for modern machine learning and ASR (Automatic Speech Recognition) workflows.
This dataset consolidates audio–text pairs from multiple open sources, applies a rigorous cleaning and normalization pipeline, and stores everything efficiently in Parquet… See the full description on the dataset page: https://huggingface.co/datasets/kiarashQ/farsi-asr-unified-cleaned.VoxPopuli-Cleaned-AA
VoxPopuli-Cleaned-AA
Quick links: AA Speech to Text Leaderboard | AA-WER v2.0 article
VoxPopuli-Cleaned-AA is a cleaned subset of the English VoxPopuli test data from esb/datasets, a speech dataset derived from European Parliament recordings. This cleaned subset is the VoxPopuli portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation of Speech to Text (STT) models.
This dataset is part of AA-WER… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/VoxPopuli-Cleaned-AA.audiosnippets-cleaned
Dataset Summary
This dataset is a processed version of mitermix/audiosnippets. The dataset contains audio snippets that have been cleaned and resampled, making it suitable for tasks like audio captioning, audio classification, or other audio-based machine learning applications.
Processing Details
Transcriptions and broken characters were removed.
All MP3 audio files were resampled to 16kHz for consistency.
The accompanying JSON metadata was made consistent.
Entries with… See the full description on the dataset page: https://huggingface.co/datasets/mkrausio/audiosnippets-cleaned.arabic-egy-cleaned
Egyptian Arabic ASR Clean 72 h
Dataset Summary
This corpus contains ≈72 h of Egyptian‑Arabic speech aligned to text.Audio has been resampled to 16 kHz mono WAV, transcripts are normalised Arabic (no diacritics, Tatweel, digits verbalised), and the data are split 80 / 10 / 10 into train / validation / test.
Supported Tasks and Leaderboards
Task
Tags
Notes
Automatic Speech Recognition
asr, speech-recognition
Primary use‑case
Forced Alignment / VAD… See the full description on the dataset page: https://huggingface.co/datasets/MAdel121/arabic-egy-cleaned.CapTTS-SFT-ears-cleanedEarnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.indicV-cleanedTH-Speech-Cleanedcommon_voice_17_t_cleanedVoxPopuli-Cleaned-AA-with-audioSame as ArtificialAnalysis/VoxPopuli-Cleaned-AA but added back the audio from the original voxpopuli dataset
ascend_MIXED_cleaned_vadascend_ZH_cleaned_vadascend_EN_cleaned_vadopen-asr-leaderboard-cleanedASR-Tamil-cleaned
Dataset Card for Dataset Name
Dataset Details
Dataset Description
This dataset is a combination of the Common Voice 16.0 and Open SLR datasets which is of 534 hours. It has been meticulously curated, normalized to a 16kHz sampling rate, and cleaned for better usability. This dataset aims to provide a comprehensive collection of speech data for various applications, including speech recognition, natural language processing, and machine learning research.… See the full description on the dataset page: https://huggingface.co/datasets/Prajwal-143/ASR-Tamil-cleaned.Common-Voice-Geo-Cleaned
Common Voice Georgian — Cleaned for TTS/STT
A high-quality subset of Mozilla Common Voice Georgian cleaned and filtered specifically for text-to-speech fine-tuning.
Dataset Summary
Total samples
21,421
Total duration
35.0 hours
Speakers
12
Sample rate
24 kHz mono WAV
Language
Georgian (kat)
Source
Mozilla Common Voice 19.0
License
CC-0 (public domain)
Splits
Split
Samples
Description
train
20,300
Training data
eval… See the full description on the dataset page: https://huggingface.co/datasets/NMikka/Common-Voice-Geo-Cleaned.tunisian-asr-cleanedluganda-english-cleaned-v1-splitbam-asr-cleaned
Origin : This is a ready-for-training clean bambara dataset good for STT models
The RobotMali's bambara asr dataset (the most recent one) especially the bambara-asr-all subset has been cleaned and converted to 16 KHz freq. to avoid overfiting during training
The the french translation column has be removed and too shorts (duration<0.9 sec) and too longs (duration > 30 secs) 's lines has been deleted too for bambara speech to text models training
Now there are the datacard will… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/bam-asr-cleaned.common_voice_17_0-cleaned_trainsong_dataset_training_20s_cleanedcommon_voice_17_0-cleanedkin_cleaned_commonvoiceacl-voice-cloning-fr-cleaned-v2
ACL Voice Cloning FR — Cleaned & Expanded (V2)
Source filtered with GPU-accelerated quality checks (SNR≥10.0dB, silence≤65%),
then expanded into all (reference, target) pairs per speaker.
Filter
Threshold
Duration
1.0-20.0s
SNR
≥ 10.0 dB
Silence
≤ 65%
Text length
5-500 chars
Split
Source kept
Expanded pairs
Shards
train
748
70,006
43
test
119
7,022
7
oromo_cleaned_datasetAlgerian-STT-Cleaned-V5cleaned_mixed_cantonese_and_english_speechAll right reserved by and credit to AlienKevin/mixed_cantonese_and_english_speech
This is a cleaned verison from AlienKevin/mixed_cantonese_and_english_speech:
https://huggingface.co/datasets/AlienKevin/mixed_cantonese_and_english_speech
Removed '"' in the preffix and suffix
Removed empty records in order to reduce hallucination
------2025-08-03------
Converted to MP3, reduce size to 1/10 of the original size.
common_voice_17_0_arabic_cleaned
