CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nader39 /audio-filesaudion<1K0 likes8.1k downloads1d agoHugging Face02D4nt3 /esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset. def add_duration(sample): y, sr = sample['audio']["array"], sample['audio']["sampling_rate"] sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000 return sample tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True) # compute duration to filter tedlium = tedlium.map(add_duration) tedlium = tedlium.select(range(512)) # Whisper max supported duration tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.audion<1K0 likes8.1k downloads2y agoHugging Face03parler-tts /libritts_r_filtered Dataset Card for Filtered LibriTTS-R This is a filtered version of LibriTTS-R. It has been filtered based on two sources: LibriTTS-R paper [1], which lists samples for which speech restoration have failed LibriTTS-P [2] list of excluded speakers for which multiple speakers have been detected. LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_filtered.audiotext-to-speech100K<n<1M24 likes4.6k downloads2y agoHugging Face04nader33 /audio-filesaudion<1K0 likes1.9k downloads2d agoHugging Face05PHBJT /cml-tts-filtered Dataset Card for Filtred and CML-TTS This dataset is a filtred version of a CML-TTS [1]. CML-TTS [1] CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in… See the full description on the dataset page: https://huggingface.co/datasets/PHBJT/cml-tts-filtered.audiotext-to-speech1M<n<10M4 likes1.6k downloads2y agoHugging Face06xhowold /file-storage-7485 File Storage Dataset This dataset is used for file storage purposes. Files This dataset contains uploaded files organized in the uploads directory. audio1K<n<10K1 likes1.6k downloads2h agoHugging Face07faker-w /ubuntu_osworld_file_cacheaudion<1K0 likes1.4k downloads3mo agoHugging Face08Eventual-Inc /sample-filesaudion<1K0 likes1.3k downloads1y agoHugging Face09nadernader /audio-filesaudio1K<n<10K0 likes632 downloads2d agoHugging Face10princed-turing /file_dataaudion<1K0 likes628 downloads3mo agoHugging Face11sapinsapin /filipinospeechcorpus Filipino Speech Corpus (FSC) Studio-recorded Filipino read, spontaneous, and word-level speech — 125 speakers, packaged as ready-to-stream Parquet. 313,322 transcribed segments · 65.1 hours · 125 speakers · 16kHz mono This is the Filipino Speech Corpus (Sagum), recorded in a controlled setting and hand/machine transcribed with Transcriber. This repo repackages the original .wav + .trs volumes as segment-level Parquet with inline audio, so you can stream it without… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/filipinospeechcorpus.audioautomatic-speech-recognition100K<n<1M3 likes609 downloads1mo agoHugging Face12MohammadGholizadeh /filimo-farsiMake Sure to use this command before downloading the dataset. !pip install "fsspec<=2023.5.0" from datasets import load_dataset import os # Define a path on your large disk for the cache cache_path = "/content/huggingface_cache" os.makedirs(cache_path, exist_ok=True) # Use the cache_dir argument to point to your new path ds = load_dataset( "MohammadGholizadeh/filimo-farsi", cache_dir=cache_path ) print(f"✅ Dataset downloaded and cached in: {cache_path}") audioautomatic-speech-recognition100K<n<1M10 likes597 downloads1y agoHugging Face13yuriyvnv /capes_synthetic_audio_filteredaudio10K<n<100K0 likes569 downloads1y agoHugging Face14bilalt1 /osworld_tasks_filesaudion<1K0 likes491 downloads7mo agoHugging Face15KaanAydinli /tsc-tr-filtered-94h-clean TSC-TR Filtered 94h — repaired transcripts ~94 hours / 72,245 utterances of Turkish TV and talk-program speech (16 kHz mono WAV) with systematically repaired transcripts. This is a derivative of ulaspolat/tsc-tr-filtered-94h, itself a filtered subset of the ISSAI Turkish Speech Corpus (MIT license). Audio is unchanged; only the text column was modified. Transcript repairs The source transcripts carry two systematic artifacts from İ/apostrophe mishandling upstream:… See the full description on the dataset page: https://huggingface.co/datasets/KaanAydinli/tsc-tr-filtered-94h-clean.audioautomatic-speech-recognition10K<n<100K0 likes477 downloads1mo agoHugging Face16qwerttyuiiop /FilSwitch FilSwitch: A Filipino–English Code-Switched Speech Dataset A Filipino-English (Taglish) read-speech dataset created for training and evaluating automatic speech recognition (ASR) systems on code-switched speech. The corpus contains 3,555 manually validated utterances from 152 speakers, totaling approximately 8.9 hours of audio. Languages Filipino + English code switching Domain Read, Philippine news Utterances 3,555 (16 kHz sampling rate) Train / Test 6.89 h /… See the full description on the dataset page: https://huggingface.co/datasets/qwerttyuiiop/FilSwitch.audioautomatic-speech-recognition1K<n<10K0 likes449 downloads1mo agoHugging Face17DavronSherbaev /uzbekvoice-filteredThis is heavy filtered version of the dataset with additional information. This dataset does not contain original Mozilla Common Voice audios or texts We filtered the dataset using number approaches: VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy Automatic STT validation. We trained the model using subset of valid… See the full description on the dataset page: https://huggingface.co/datasets/DavronSherbaev/uzbekvoice-filtered.audioautomatic-speech-recognition100K<n<1M19 likes421 downloads2y agoHugging Face18kiriku /Homophones_filted_datasetaudio100K<n<1M0 likes330 downloads3y agoHugging Face19FILM6912 /th-en-zh-tts-200k-enhanced TH-EN-ZH Multi-speaker TTS Dataset (200K, RE-USE Enhanced) Speech-enhanced variant of FILM6912/th-en-zh-tts-200k. Every clip has been processed through NVIDIA RE-USE (universal speech enhancement, SEMamba) at its native sample rate, then re-encoded losslessly as FLAC (PCM_16). Same schema, same row order, same 200,000 rows (th 100k / en 50k / zh 50k): Column Type Description text string Transcript (identical to the original dataset) audio Audio Enhanced audio, FLAC… See the full description on the dataset page: https://huggingface.co/datasets/FILM6912/th-en-zh-tts-200k-enhanced.audiotext-to-speech100K<n<1M0 likes321 downloads2d agoHugging Face20vsrinivas /bengali_audio_files Dataset Card for "bengali_audio_files" More Information needed audio100K<n<1M0 likes318 downloads3y agoHugging Face21ghananlpcommunity /filtered_ghana_asr This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Filtered Ghana Asr audio100K<n<1M0 likes301 downloads3mo agoHugging Face22SilencioNetwork /tagalog-filipino-speech Tagalog / Filipino Spontaneous Speech — Silencio Philippines Pack Spontaneous Tagalog/Filipino with human transcription and word-level forced alignment. Sixteen speakers, 90 unscripted clips, 13,782 timestamped tokens. Part of the Silencio Philippines Pack. Hours 2.03 Clips 90 Speakers 16 Countries 1 Speaker origin regions 2 L1 speakers of the recorded language 15 of 16 (86 clips) Audio 48 kHz stereo WAV Mean clip length 81.0 s Transcripts… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/tagalog-filipino-speech.audioautomatic-speech-recognitionn<1K1 likes292 downloads2d agoHugging Face23marie998 /files_2audion<1K1 likes274 downloads1y agoHugging Face24hamsaai /Recorrected_Classification_Data_filtered_trainaudio10K<n<100K0 likes256 downloads2mo agoHugging Face25mrfakename /noisy-speech-filesaudio1 likes214 downloads2y agoHugging Face26ulaspolat /tsc-tr-filtered-94h Dataset Card for Turkish Speech Corpus (TSC) — Preprocessed Edition Dataset Summary This dataset is a preprocessed and filtered version of the Turkish Speech Corpus (TSC), originally published by the Institute of Smart Systems and Artificial Intelligence (ISSAI) at Nazarbayev University. The original corpus contains 218.2 hours of transcribed Turkish speech across 186,171 utterances and is described in the paper Multilingual Speech Recognition for Turkic Languages… See the full description on the dataset page: https://huggingface.co/datasets/ulaspolat/tsc-tr-filtered-94h.audioautomatic-speech-recognition10K<n<100K0 likes214 downloads3mo agoHugging Face27nkudratov /filtered_uzbekvoiceaudio100K<n<1M0 likes210 downloads3mo agoHugging Face28ittailup /filtered_common_voice-enaudio100K<n<1M0 likes204 downloads2y agoHugging Face29ai4uz /uzbekvoice-filteredThis is heavy filtered version of the dataset with additional information. This dataset does not contain original Mozilla Common Voice audios or texts We filtered the dataset using number approaches: VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy Automatic STT validation. We trained the model using subset of valid… See the full description on the dataset page: https://huggingface.co/datasets/ai4uz/uzbekvoice-filtered.audioautomatic-speech-recognition100K<n<1M0 likes194 downloads5mo agoHugging Face30krishnakalyan3 /emo_speech_filtered_v12 second filtered emotional speech in webdataset format https://huggingface.co/datasets/EQ4You/Emotional_Speech audio10K<n<100K0 likes181 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.