datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audio-filesesb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset.
def add_duration(sample):
y, sr = sample['audio']["array"], sample['audio']["sampling_rate"]
sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000
return sample
tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True)
# compute duration to filter
tedlium = tedlium.map(add_duration)
tedlium = tedlium.select(range(512))
# Whisper max supported duration
tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.libritts_r_filtered
Dataset Card for Filtered LibriTTS-R
This is a filtered version of LibriTTS-R. It has been filtered based on two sources:
LibriTTS-R paper [1], which lists samples for which speech restoration have failed
LibriTTS-P [2] list of excluded speakers for which multiple speakers have been detected.
LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately
585 hours of read English speech at 24kHz sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_filtered.audio-filescml-tts-filtered
Dataset Card for Filtred and CML-TTS
This dataset is a filtred version of a CML-TTS [1].
CML-TTS [1] CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in… See the full description on the dataset page: https://huggingface.co/datasets/PHBJT/cml-tts-filtered.file-storage-7485
File Storage Dataset
This dataset is used for file storage purposes.
Files
This dataset contains uploaded files organized in the uploads directory.
ubuntu_osworld_file_cachesample-filesaudio-filesfile_datafilipinospeechcorpus
Filipino Speech Corpus (FSC)
Studio-recorded Filipino read, spontaneous, and word-level speech — 125 speakers, packaged as ready-to-stream Parquet.
313,322 transcribed segments · 65.1 hours · 125 speakers · 16kHz mono
This is the Filipino Speech Corpus (Sagum), recorded in a controlled setting and
hand/machine transcribed with Transcriber. This repo
repackages the original .wav + .trs volumes as segment-level Parquet with
inline audio, so you can stream it without… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/filipinospeechcorpus.filimo-farsiMake Sure to use this command before downloading the dataset.
!pip install "fsspec<=2023.5.0"
from datasets import load_dataset
import os
# Define a path on your large disk for the cache
cache_path = "/content/huggingface_cache"
os.makedirs(cache_path, exist_ok=True)
# Use the cache_dir argument to point to your new path
ds = load_dataset(
"MohammadGholizadeh/filimo-farsi",
cache_dir=cache_path
)
print(f"✅ Dataset downloaded and cached in: {cache_path}")
capes_synthetic_audio_filteredosworld_tasks_filestsc-tr-filtered-94h-clean
TSC-TR Filtered 94h — repaired transcripts
~94 hours / 72,245 utterances of Turkish TV and talk-program speech (16 kHz
mono WAV) with systematically repaired transcripts. This is a derivative of
ulaspolat/tsc-tr-filtered-94h,
itself a filtered subset of the ISSAI Turkish Speech Corpus
(MIT license). Audio is unchanged; only the text column was modified.
Transcript repairs
The source transcripts carry two systematic artifacts from İ/apostrophe
mishandling upstream:… See the full description on the dataset page: https://huggingface.co/datasets/KaanAydinli/tsc-tr-filtered-94h-clean.FilSwitch
FilSwitch: A Filipino–English Code-Switched Speech Dataset
A Filipino-English (Taglish) read-speech dataset created for training and evaluating automatic speech recognition (ASR) systems on code-switched speech. The corpus contains 3,555 manually validated utterances from 152 speakers, totaling approximately 8.9 hours of audio.
Languages
Filipino + English code switching
Domain
Read, Philippine news
Utterances
3,555 (16 kHz sampling rate)
Train / Test
6.89 h /… See the full description on the dataset page: https://huggingface.co/datasets/qwerttyuiiop/FilSwitch.uzbekvoice-filteredThis is heavy filtered version of the dataset with additional information.
This dataset does not contain original Mozilla Common Voice audios or texts
We filtered the dataset using number approaches:
VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed
Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy
Automatic STT validation. We trained the model using subset of valid… See the full description on the dataset page: https://huggingface.co/datasets/DavronSherbaev/uzbekvoice-filtered.Homophones_filted_datasetth-en-zh-tts-200k-enhanced
TH-EN-ZH Multi-speaker TTS Dataset (200K, RE-USE Enhanced)
Speech-enhanced variant of FILM6912/th-en-zh-tts-200k.
Every clip has been processed through NVIDIA RE-USE (universal speech enhancement, SEMamba) at its native sample rate, then re-encoded losslessly as FLAC (PCM_16).
Same schema, same row order, same 200,000 rows (th 100k / en 50k / zh 50k):
Column
Type
Description
text
string
Transcript (identical to the original dataset)
audio
Audio
Enhanced audio, FLAC… See the full description on the dataset page: https://huggingface.co/datasets/FILM6912/th-en-zh-tts-200k-enhanced.bengali_audio_files
Dataset Card for "bengali_audio_files"
More Information needed
filtered_ghana_asr
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Filtered Ghana Asr
tagalog-filipino-speech
Tagalog / Filipino Spontaneous Speech — Silencio Philippines Pack
Spontaneous Tagalog/Filipino with human transcription and word-level forced alignment. Sixteen speakers, 90 unscripted clips, 13,782 timestamped tokens. Part of the Silencio Philippines Pack.
Hours
2.03
Clips
90
Speakers
16
Countries
1
Speaker origin regions
2
L1 speakers of the recorded language
15 of 16 (86 clips)
Audio
48 kHz stereo WAV
Mean clip length
81.0 s
Transcripts… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/tagalog-filipino-speech.files_2Recorrected_Classification_Data_filtered_trainnoisy-speech-filestsc-tr-filtered-94h
Dataset Card for Turkish Speech Corpus (TSC) — Preprocessed Edition
Dataset Summary
This dataset is a preprocessed and filtered version of the Turkish Speech Corpus (TSC), originally published by the Institute of Smart Systems and Artificial Intelligence (ISSAI) at Nazarbayev University. The original corpus contains 218.2 hours of transcribed Turkish speech across 186,171 utterances and is described in the paper Multilingual Speech Recognition for Turkic Languages… See the full description on the dataset page: https://huggingface.co/datasets/ulaspolat/tsc-tr-filtered-94h.filtered_uzbekvoicefiltered_common_voice-enuzbekvoice-filteredThis is heavy filtered version of the dataset with additional information.
This dataset does not contain original Mozilla Common Voice audios or texts
We filtered the dataset using number approaches:
VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed
Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy
Automatic STT validation. We trained the model using subset of valid… See the full description on the dataset page: https://huggingface.co/datasets/ai4uz/uzbekvoice-filtered.emo_speech_filtered_v12 second filtered emotional speech in webdataset format
https://huggingface.co/datasets/EQ4You/Emotional_Speech
