datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
libritts_r_filtered
Dataset Card for Filtered LibriTTS-R
This is a filtered version of LibriTTS-R. It has been filtered based on two sources:
LibriTTS-R paper [1], which lists samples for which speech restoration have failed
LibriTTS-P [2] list of excluded speakers for which multiple speakers have been detected.
LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately
585 hours of read English speech at 24kHz sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_filtered.osworld_tasks_filesfilipinospeechcorpus
Filipino Speech Corpus (FSC)
Studio-recorded Filipino read, spontaneous, and word-level speech — 125 speakers, packaged as ready-to-stream Parquet.
313,322 transcribed segments · 65.1 hours · 125 speakers · 16kHz mono
This is the Filipino Speech Corpus (Sagum), recorded in a controlled setting and
hand/machine transcribed with Transcriber. This repo
repackages the original .wav + .trs volumes as segment-level Parquet with
inline audio, so you can stream it without… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/filipinospeechcorpus.filimo-farsiMake Sure to use this command before downloading the dataset.
!pip install "fsspec<=2023.5.0"
from datasets import load_dataset
import os
# Define a path on your large disk for the cache
cache_path = "/content/huggingface_cache"
os.makedirs(cache_path, exist_ok=True)
# Use the cache_dir argument to point to your new path
ds = load_dataset(
"MohammadGholizadeh/filimo-farsi",
cache_dir=cache_path
)
print(f"✅ Dataset downloaded and cached in: {cache_path}")
tsc-tr-filtered-94h-clean
TSC-TR Filtered 94h — repaired transcripts
~94 hours / 72,245 utterances of Turkish TV and talk-program speech (16 kHz
mono WAV) with systematically repaired transcripts. This is a derivative of
ulaspolat/tsc-tr-filtered-94h,
itself a filtered subset of the ISSAI Turkish Speech Corpus
(MIT license). Audio is unchanged; only the text column was modified.
Transcript repairs
The source transcripts carry two systematic artifacts from İ/apostrophe
mishandling upstream:… See the full description on the dataset page: https://huggingface.co/datasets/KaanAydinli/tsc-tr-filtered-94h-clean.FilSwitch
FilSwitch: A Filipino–English Code-Switched Speech Dataset
A Filipino-English (Taglish) read-speech dataset created for training and evaluating automatic speech recognition (ASR) systems on code-switched speech. The corpus contains 3,555 manually validated utterances from 152 speakers, totaling approximately 8.9 hours of audio.
Languages
Filipino + English code switching
Domain
Read, Philippine news
Utterances
3,555 (16 kHz sampling rate)
Train / Test
6.89 h /… See the full description on the dataset page: https://huggingface.co/datasets/qwerttyuiiop/FilSwitch.uzbekvoice-filteredThis is heavy filtered version of the dataset with additional information.
This dataset does not contain original Mozilla Common Voice audios or texts
We filtered the dataset using number approaches:
VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed
Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy
Automatic STT validation. We trained the model using subset of valid… See the full description on the dataset page: https://huggingface.co/datasets/DavronSherbaev/uzbekvoice-filtered.filimo-persian-asrThis dataset consists of about 400 hours of audio extracted from various Filimo videos in the Persian language.
Note: This dataset contains raw, unvalidated transcriptions. Users are advised to:
1. Perform their own quality assessment
2. Create their own train/validation/test splits based on their specific needs
3. Validate a subset of the data if needed for their use casetagalog-filipino-speech
Tagalog / Filipino Spontaneous Speech — Silencio Philippines Pack
Spontaneous Tagalog/Filipino with human transcription and word-level forced alignment. Sixteen speakers, 90 unscripted clips, 13,782 timestamped tokens. Part of the Silencio Philippines Pack.
Hours
2.03
Clips
90
Speakers
16
Countries
1
Speaker origin regions
2
L1 speakers of the recorded language
15 of 16 (86 clips)
Audio
48 kHz stereo WAV
Mean clip length
81.0 s
Transcripts… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/tagalog-filipino-speech.tsc-tr-filtered-94h
Dataset Card for Turkish Speech Corpus (TSC) — Preprocessed Edition
Dataset Summary
This dataset is a preprocessed and filtered version of the Turkish Speech Corpus (TSC), originally published by the Institute of Smart Systems and Artificial Intelligence (ISSAI) at Nazarbayev University. The original corpus contains 218.2 hours of transcribed Turkish speech across 186,171 utterances and is described in the paper Multilingual Speech Recognition for Turkic Languages… See the full description on the dataset page: https://huggingface.co/datasets/ulaspolat/tsc-tr-filtered-94h.uzbekvoice-filteredThis is heavy filtered version of the dataset with additional information.
This dataset does not contain original Mozilla Common Voice audios or texts
We filtered the dataset using number approaches:
VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed
Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy
Automatic STT validation. We trained the model using subset of valid… See the full description on the dataset page: https://huggingface.co/datasets/ai4uz/uzbekvoice-filtered.no-filter-raw-NepaliParliamentDSv2Eng-Filipino-Accented-audio-with-human-transcription-call-center-topicThis dataset contains 103+ hours of spontaneous English conversations spoken in a Filipino accent, recorded in a studio environment to ensure crystal-clear audio quality. The conversations are designed as role-play scenarios between agents and customers across a variety of call center domains.
🗣️ Speech Style: Natural, unscripted role-playing between native Filipino-accented English speakers, simulating real-world customer interactions.
🎧 Audio Format: High-quality stereo WAV files, recorded… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Eng-Filipino-Accented-audio-with-human-transcription-call-center-topic.ascend
Dataset Card for ASCEND
Dataset Summary
ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/filwsyl/ascend.librivox_filtered_id
Librivox Filtered ID
Filtered Librivox Indonesian dataset
Audio has been preprocessed using FFmpeg as: wav -ar 16000 -ac 1 (mono 16kHz sample_rate) for Whisper-ready finetuning
Selected audio datasets on: ['id']['universal-declaration-of-human-rights']
num_rows: 136
Original dataset: indonesian-nlp/librivox-indonesia
Format
Each example is a dictionary with the following fields:
{
"path": "audio/librivox_id_1.wav",
"audio": {
"path": "audio/librivox_id_1.wav"… See the full description on the dataset page: https://huggingface.co/datasets/Willy030125/librivox_filtered_id.Filipino-Tagalog-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 169 hours of processed Filipino (FIL) and 4019 hours of processed Tagalog (TL) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Filipino-Tagalog-Call-Center-Audio-Dataset-Single-Channel.Filipino_Tagalog_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 169 hours of processed Filipino (FIL) and 4,019 hours of processed Tagalog (TL) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Filipino_Tagalog_Call_Center_Audio_Dataset_Dual_Channel.tibetan-audio-to-english-fixed-filtered
Tibetan audio translation Dataset
Dataset Description
Tibetan audio translation Dataset
Dataset Summary
This dataset contains 6,366 audio samples with corresponding transcriptions, totaling approximately 15.8 hours of audio.
Languages
The dataset is in EN (Language code: en).
Dataset Structure
Data Fields
audio: An audio object containing:
path: Path to the audio file (if applicable)
array: Audio waveform as a numpy array… See the full description on the dataset page: https://huggingface.co/datasets/Titung/tibetan-audio-to-english-fixed-filtered.eka_filtered_data
Eka Medical Asr Sample Noise Eval Dataset
Dataset Description
This dataset contains 62 samples organized across multiple splits and 31 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
original: 2 samples
test: 2 samples
noisy-bg-snr-10: 2 samples
test: 2 samples
noisy-bg-snr-20: 2 samples
test: 2 samples
noisy-bg-snr-30: 2 samples
test: 2 samples
noisy-bg-snr-40: 2 samples
test:… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/eka_filtered_data.uzbekvoice-filtered2This is heavy filtered version of the dataset with additional information.
This dataset does not contain original Mozilla Common Voice audios or texts
We filtered the dataset using number approaches:
VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed
Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy
Automatic STT validation. We trained the model using subset of valid… See the full description on the dataset page: https://huggingface.co/datasets/Jurabek/uzbekvoice-filtered2.gdrive-src0506-pesq240-filtered
gdrive src05 + src06 — PESQ > 2.4 filtered subset
395 Yoruba speech clips selected from two internal audio-chunk corpora by
BiCodec reconstruction quality, shipped with transcripts and with the codec
reconstruction of each clip alongside the original.
Yield
Source
Scored
Kept (PESQ > 2.4)
Rate
gdrive_src05
1000
158
15.80%
gdrive_src06
997
237
23.77%
merged
1997
395
19.78%
Languages — read this before filtering
The two sources are… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-src0506-pesq240-filtered.acp-corpus-filtered
ACP Filtered Conversations (via Cortico)
This dataset is a filtered slice of conversation recordings and transcripts
from the American Conversation Project (ACP), retrieved via
Cortico's ACP integration. "Filtered" means every
fragment included here already passed an LLM-based salience pass (the
project's internal "wheat vs. chaff" filter) that removed filler, small talk,
and interjections — everything kept is a substantive, quote-anchored moment
someone actually said.
This is… See the full description on the dataset page: https://huggingface.co/datasets/omzugo/acp-corpus-filtered.FILALIHicham_CommonVoice-whisper-processed-newCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données FILALIHicham/CommonVoice-whisper-processed-new.
filtered-gol-dataset
Filtered GOL Dataset
midralab/gol-dataset をTTS(Text-to-Speech)学習用にフィルタリングしたデータセットです。
データセット概要
項目
値
総再生時間
約1,880時間
サンプル数
約120万
話者数
380人
データサイズ
約280GB
形式
WebDataset (.tar)
音声形式
FLAC (44.1kHz, モノラル)
フィルタリング条件
基本フィルタ
テキスト長: 3文字以上
音声長: 1秒以上、60秒未満
話者フィルタ
話者あたり5時間以上の音声データを持つ話者のみ
テキストフィルタ(除外対象)
非言語テキスト(句読点のみ、空白のみなど)
顔文字 (^_^), (T_T) など
笑い表現 (笑), 文末の www
絵文字
英数字のみのテキスト
同一文字4回以上の繰り返し
データ構造… See the full description on the dataset page: https://huggingface.co/datasets/tts-dataset/filtered-gol-dataset.reazonspeech_qwen3-asr_large_filtered
Summary
This is the ReazonSpeech corpus's large split, featuring Qwen3-ASR 1.7B transcriptions.
Since the original transcriptions often contain errors, comparing them with the Qwen3-ASR outputs could be useful.
This is a filtered version in which no insertions occurred from the original transcriptions to the Qwen3-ASR transcriptions.
Usage
import json
import webdataset as wds
from huggingface_hub import get_token
SHARDS = (
f"pipe:curl -sLf -H 'Authorization: Bearer… See the full description on the dataset page: https://huggingface.co/datasets/nagohachi/reazonspeech_qwen3-asr_large_filtered.FILALIHicham_LibriSpeech-whisper-processed-newCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données FILALIHicham/LibriSpeech-whisper-processed-new.
EN-MALAY-CS-FILTERED
EN-MALAY-CS-FILTERED
English–Malay code-switching conversational speech from the IMDA National
Speech Corpus (2021), segmented to utterance level and cleaned with a
3-model agreement filter: each utterance was transcribed by three ASR
models — openai/whisper-large-v3, MERaLiON/MERaLiON-2-10B-ASR, and
Qwen/Qwen3-ASR-1.7B — and an utterance is removed when all three
models score WER > 60% against the reference transcript (all models
agreeing the reference is unreliable).… See the full description on the dataset page: https://huggingface.co/datasets/yyhenggg/EN-MALAY-CS-FILTERED.FILALIHicham_TedX-whisper-processed-newCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données FILALIHicham/TedX-whisper-processed-new.
FILALIHicham_Ofrom-whisper-processed-newCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données FILALIHicham/Ofrom-whisper-processed-new.
FILALIHicham_Mediaspeech-whisper-processedCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données FILALIHicham/Mediaspeech-whisper-processed.
