CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01no7z /hsk-sentences-audio HSK Sentences Audio 4,354 Chinese sentences graded against the official HSK 3.0 levels 1–6, with pinyin, English translations, per-word glosses, grammar tags, and normal/slow synthetic speech. The complete export contains 8,708 MP3 files. Dataset structure The Viewer reads native Parquet from data/train.parquet, avoiding a dependency on Hugging Face's JSON-to-Parquet conversion service. The same 4,354 records are also available as validated JSON Lines in… See the full description on the dataset page: https://huggingface.co/datasets/no7z/hsk-sentences-audio.audiotext-to-speech1K<n<10K0 likes648 downloads2mo agoHugging Face02TCabbage /gsat-vocab-sentences-tts GSAT Vocabulary TTS Audio Text-to-speech audio files for GSAT (General Scholastic Ability Test) English vocabulary. Structure audio/ - MP3 audio files organized by hash prefix (e.g., audio/ab/abcd1234....mp3) index.jsonl - Index file mapping hashes to text and TTS engine used Engines Kokoro (af_heart voice) - Used for lemmas (single words/phrases) Supertonic (M1 voice) - Used for example sentences Audio Format Format: MP3 Sample rate: 24kHz… See the full description on the dataset page: https://huggingface.co/datasets/TCabbage/gsat-vocab-sentences-tts.audiotext-to-speech10K<n<100K0 likes322 downloads8mo agoHugging Face03sarahwei /Taiwanese-Minnan-Example-Sentences Taiwanese Minnan Example Sentences The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems. Dataset Features Source: Ministry of Education, Taiwan (Sutian Resource Center) Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.audioautomatic-speech-recognition10K<n<100K12 likes301 downloads2y agoHugging Face04TartarusXXX /mixed-language-detection-pilot-complete-sentences Mixed-Language Speech Detection Pilot — Complete Sentences This is the complete-sentence revision of a 6,000-clip binary audio-classification pilot. label = 0 denotes one intended language and label = 1 denotes more than one intended language. The covered languages are Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani (ckb), Arabic (ara), Persian (fas), and English (eng). What changed Earlier generation forced source transcripts into arbitrary… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-complete-sentences.audioaudio-classification1K<n<10K0 likes191 downloads1mo agoHugging Face05Reza2kn /visualears-hardword-sentences 🗂️ visualears-hardword-sentences English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Hard-word sentence dataset used for semantic/keyword stress cases. جمله‌های دارای واژه‌های دشوار و معنایی برای آزمون تنش واژگانی، بازیابی کلیدواژه و S³. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 266 files; approximately 39.23 GB 266 فایل؛ حدود 39.23… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-hardword-sentences.audio100K<n<1M2 likes180 downloads2mo agoHugging Face06danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes152 downloads10mo agoHugging Face07blanchon /ears_dataset_sentencesaudio10K<n<100K3 likes104 downloads2y agoHugging Face08i4ds /swiss-german-city-sentences_trainaudio1K<n<10K2 likes100 downloads5mo agoHugging Face09i4ds /swiss-german-city-sentences_valaudion<1K1 likes91 downloads5mo agoHugging Face10i4ds /swiss-german-city-sentences_v2 Swiss German City Sentences v2 Synthetic Swiss German speech dataset with city name sentences across multiple dialects. audioautomatic-speech-recognition10K<n<100K1 likes76 downloads4mo agoHugging Face11danielrosehill /English-Hebrew-Mixed-Sentences English-Hebrew Mixed Sentences Dataset A dataset of English sentences with Hebrew words and phrases interspersed, designed for speech-to-text training and evaluation for English speakers in Israel. Overview This dataset addresses a common challenge for English-speaking immigrants in Israel: standard speech-to-text (STT) systems struggle to accurately transcribe code-switched speech where Hebrew words are mixed into primarily English sentences. Example: "I need to pick up… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/English-Hebrew-Mixed-Sentences.audion<1K0 likes63 downloads10mo agoHugging Face12transitionGap /ASR_Marathi_Sentences Marathi Sentence-Level ASR Dataset 📌 Overview This dataset contains sentence-level Marathi speech segments aligned with transcripts. The dataset was created by extracting subtitle timestamps (SRV3 format) from Marathi YouTube content and segmenting the corresponding audio using precise time alignment. Each sample contains: A WAV audio file (sentence-level) The corresponding Marathi transcript text This dataset is suitable for: Whisper fine-tuning Wav2Vec2 CTC training… See the full description on the dataset page: https://huggingface.co/datasets/transitionGap/ASR_Marathi_Sentences.audio1K<n<10K0 likes37 downloads7mo agoHugging Face13danny666 /child_handpicked_sentencesaudion<1K0 likes37 downloads1mo agoHugging Face14transitionGap /ASR_Hanyanvi_7k_sentencesaudio1K<n<10K0 likes27 downloads7mo agoHugging Face15transitionGap /ASR_Haryanvi_1K_sentencesaudio1K<n<10K0 likes22 downloads7mo agoHugging Face16mia-project /child_handpicked_sentencesaudion<1K0 likes16 downloads3y agoHugging Face17transitionGap /ASR_1kBhojpuri_Sentencesaudio1K<n<10K0 likes15 downloads7mo agoHugging Face18aomocelin /synthetic_utterances_pt-BR_1400_sentences_x_6_speakersaudio1K<n<10K0 likes13 downloads3mo agoHugging Face19foxhound /camara_audio_sentencesaudio10K<n<100K0 likes13 downloads2mo agoHugging Face20transitionGap /ASR_1kHindi_Sentencesaudio1K<n<10K0 likes12 downloads7mo agoHugging Face21aomocelin /synthetic_utterances_pt-BR_1200_sentences_x_6_speakersaudio1K<n<10K0 likes12 downloads3mo agoHugging Face22deepinfinityai /30_report_sentences_datasetaudion<1K0 likes6 downloads2y agoHugging Face23adithyal1998Bhat /tts_synthetic_kn_single_sentencesaudio1K<n<10K0 likes5 downloads1y agoHugging Face24Prasad12344321 /ASR_Marathi_Sentences Marathi Sentence-Level ASR Dataset 📌 Overview This dataset contains sentence-level Marathi speech segments aligned with transcripts. The dataset was created by extracting subtitle timestamps (SRV3 format) from Marathi YouTube content and segmenting the corresponding audio using precise time alignment. Each sample contains: A WAV audio file (sentence-level) The corresponding Marathi transcript text This dataset is suitable for: Whisper fine-tuning Wav2Vec2 CTC training… See the full description on the dataset page: https://huggingface.co/datasets/Prasad12344321/ASR_Marathi_Sentences.audio1K<n<10K0 likes4 downloads7mo agoHugging Face25abuelnasr /eg-ADI-sentencesgatedaudio1K<n<10K0 likes3 downloads2y agoHugging Face26adilet0000 /harvard-sentences-kokoroaudion<1K0 likes3 downloads8mo agoHugging Face27mrfluffypants /kokoro-harvard-sentencesaudion<1K0 likes3 downloads8mo agoHugging Face28Prasad12344321 /ASR_Marathi_Sentences_1kaudio1K<n<10K0 likes3 downloads7mo agoHugging Face29NgQuocThai /S2T_SplitEndMovie_Sentencesgatedaudio10K<n<100K0 likes2 downloads1y agoHugging Face30NgQuocThai /S2T_Merged_Sentencesgatedaudio10K<n<100K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.