CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes11k downloads8h agoHugging Face02espnet /Bagpiper_PreTrain_Data Bagpiper Pretraining Data Bagpiper Pretraining Data is the public rich-captioned audio snapshot associated with Bagpiper, an open-ended audio language model that learns bidirectional mappings between audio and comprehensive text descriptions across speech, music, environmental sound, and mixtures. The en metadata describes the primary rich-caption language. Source audio can contain speech or singing in other languages; it is not an English-only audio guarantee. The repository… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_PreTrain_Data.tabularautomatic-speech-recognition10K<n<100K0 likes2.6k downloads2mo agoHugging Face03anuj-inavlabs /kupe-asr-en-data kupe-asr-en-mini-150m — data Two loadable configs, packed into ~20-25 bunch_*.parquet files each (Hub-quota friendly): raw — 24 kHz mono English audio (flac bytes) + text. The encode stage reads this. mimi — Mimi c0..c7 codes (12.5 Hz) + text. Training reads this. Ledgers under ledger/ (data.json, mimi.json) track collected/encoded hours and resume state. from datasets import load_dataset ds = load_dataset("anuj-inavlabs/kupe-asr-en-data", "mimi", split="train") tabularautomatic-speech-recognition1M<n<10M0 likes1k downloads13d agoHugging Face04humair025 /hashed_data Munch Hashed Index - Lightweight Audio Reference Dataset 📖 Overview Munch Hashed Index is a lightweight reference dataset that provides SHA-256 hashes for all audio files in the Munch Urdu TTS Dataset. Instead of storing 1.27 TB of raw audio, this index stores only metadata and cryptographic hashes, enabling: ✅ Fast duplicate detection across 4.17 million audio samples ✅ Efficient dataset exploration without downloading terabytes ✅ Quick metadata queries (voice… See the full description on the dataset page: https://huggingface.co/datasets/humair025/hashed_data.tabulartext-generation1M<n<10M0 likes984 downloads10mo agoHugging Face05OmniEvalKit /omnievalkit-dataset OmniEvalKit Evaluation Datasets Evaluation datasets for OmniEvalKit, a comprehensive evaluation framework for omni-modal (audio + video + image + text) models. Overview Total subsets: 65 Total samples: 315,264 Total size: 620.3 GB (Parquet with embedded audio/image/video) Subsets with embedded video: 15 Subsets requiring external video download: 2 Usage from datasets import load_dataset ds = load_dataset("OmniEvalKit/omnievalkit-dataset", "aishell1_test")… See the full description on the dataset page: https://huggingface.co/datasets/OmniEvalKit/omnievalkit-dataset.audioaudio-classification100K<n<1M0 likes662 downloads6mo agoHugging Face06RodolfoRizzi /MyMentorLLM-dataset Dataset Card for MyMentorLLM This dataset accompanies the MyMentorLLM paper, arXiv:2607.25667, which describes the simulation environment, generation procedure, experimental design and validation analyses. If you use this dataset, please cite the paper (see Citation Information). Dataset Summary MyMentorLLM is a multimodal voice- and text-based deliberate-practice environment used to generate 2,100 simulated Cognitive Behavioural Therapy (CBT) training sessions… See the full description on the dataset page: https://huggingface.co/datasets/RodolfoRizzi/MyMentorLLM-dataset.audiotext-generation10K<n<100K0 likes452 downloads2d agoHugging Face07Atika88 /Indonesian-ASR-11-Class-Dataset Indonesian ASR 11-Class Dataset Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts. Dataset summary Audio files: 104,500 WAV files Real/human recordings: 104,368 Synthetic repair files: 132 Sentence classes: 11 Indonesian sentence categories Canonical balanced sentence slots: 209 (11 categories × 19 retained slots) Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs* Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.tabularautomatic-speech-recognition100K<n<1M0 likes324 downloads14d agoHugging Face08paodigitalhub /pao-audio-dataset 🎙️ Pa'O Audio Dataset ပအိုဝ်ႏ အငေါဝ်း အဆင်ႏဗာႏ ရွမ်ခြွဉ်းဗူႏ 📌 Project Summary The Pa'O Audio Dataset is an open-source initiative created to facilitate the development of speech technologies and Artificial Intelligence tools for the Pa'O language (ပအိုဝ်ႏဘာႏသာႏငေါဝ်းငွါ). Pa'O is primarily spoken in Shan State and other regions of Myanmar. As a low-resource language in the AI landscape, this dataset provides audio recordings and corresponding… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-audio-dataset.audioautomatic-speech-recognitionn<1K1 likes238 downloads10d agoHugging Face09FBK-MT /fama-data Dataset Description, Collection, and Source The FAMA training data is the collection of English and Italian datasets for automatic speech recognition (ASR) and speech translation (ST) used to train the FAMA models family. The ASR section of FAMA is derived from the MOSEL data collection, including the automatic transcripts obtained with Whisper and available in the HuggingFace MOSEL Dataset. The ASR is further augmented with automatically transcribed speech from the… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/fama-data.tabulartranslation1M<n<10M2 likes196 downloads1y agoHugging Face10Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes137 downloads1mo agoHugging Face11Sakchham19 /Trump_Voice_Dataset Trump Voice Dataset This dataset contains audio clips of Donald Trump's speech from the World Economic Forum (WEF) 2018, paired with their corresponding transcriptions. The dataset is designed for text-to-speech (TTS) and speech recognition tasks. Dataset Description Dataset Summary The Trump Voice Dataset consists of 20 audio samples (10 train, 10 test) extracted from Donald Trump's speech at the World Economic Forum 2018. Each audio clip is approximately 10… See the full description on the dataset page: https://huggingface.co/datasets/Sakchham19/Trump_Voice_Dataset.tabularautomatic-speech-recognitionn<1K0 likes55 downloads1y agoHugging Face12kvest /Swedia-ASR-Dataset Swedia ASR Dataset This repository contains a small Swedish ASR evaluation dataset based on speech transcriptions from Swedia 2000. It was assembled to compare automatic speech-recognition output against manually corrected reference transcriptions for Swedish dialectal speech. The dataset is useful for quick experiments with Swedish ASR systems, especially when you want to inspect recognition quality on spontaneous speech from different regions, speakers, ages, and genders.… See the full description on the dataset page: https://huggingface.co/datasets/kvest/Swedia-ASR-Dataset.tabularautomatic-speech-recognitionn<1K1 likes50 downloads5mo agoHugging Face13FatimahEmadEldin /Arabic-Emotional-Audio-Dataset-Baved BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging) A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits. Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.audioaudio-classification1K<n<10K0 likes49 downloads5mo agoHugging Face14IMJONEZZ /star-wars-dataset Star Wars dialogue dataset - annotation layers (snapshot 2026-09-21) One master cue table serving diarization, ASR and dialogue-LLM training: every subtitle cue part of 15 titles with a forced-aligned time span, a character label and provenance. No audio or video is included - the source media is copyrighted. Clip paths in the manifests resolve once you rebuild the clips from your own copies with the pipeline code (export_asr.py, export_diarization.py). The text layers contain… See the full description on the dataset page: https://huggingface.co/datasets/IMJONEZZ/star-wars-dataset.tabularautomatic-speech-recognition10K<n<100K2 likes38 downloads1d agoHugging Face15Letian2003 /stage1a_smoke_data stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh) Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP → frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend: WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT encoder — no raw-audio decoding at train time. 113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.tabularautomatic-speech-recognitionn<1K0 likes22 downloads2mo agoHugging Face16Pedramebd /welsh-speech-dataset Welsh Speech Dataset A multimodal dataset of 33 speakers producing 10 Welsh phrases, captured using 3DMD technology with audio and dense facial landmark annotations. Dataset Overview Speakers: 33 participants Phrases: 10 Welsh phrases per speaker Sequences: ~330 (33 speakers x 10 phrases) Modalities: Audio recordings (.wav) 3D facial reconstructions (.obj meshes + texture maps) 68-point facial landmarks (ibug68 template) Fluency Scores: Each phrase rated 0-5… See the full description on the dataset page: https://huggingface.co/datasets/Pedramebd/welsh-speech-dataset.tabularautomatic-speech-recognitionn<1K0 likes22 downloads1mo agoHugging Face17DrUkachi /ktt-math-tutor-data KTT Math Tutor — Data Data artefacts for the AIMS KTT Hackathon Tier-3 submission S2.T3.1 AI Math Tutor for Early Learners. Source code: https://github.com/DrUkachi/ktt-math-tutor. Contents T3.1_Math_Tutor/ Core curriculum + seeds. curriculum.json — 80 items × 5 sub-skills (counting, number sense, addition, subtraction, word problem) with EN / FR / KIN stems, difficulty 1–10, age bands 5–6 / 6–7 / 7–8 / 8–9, visual asset keys, expected integer answer.… See the full description on the dataset page: https://huggingface.co/datasets/DrUkachi/ktt-math-tutor-data.tabularquestion-answeringn<1K0 likes18 downloads5mo agoHugging Face18ayush712145 /sarvam-tts-dataset Sarvam TTS Dataset: Indian English and Hindi A curated speech dataset for TTS model training, containing 396 clips in Indian English (en-IN) and Hindi (hi-IN) totalling 57.3 minutes (3440 seconds). Pipeline source code: https://github.com/Ayush147258/sarvam-tts-dataset Dataset Summary Stat Value Total clips 396 Total duration 57.3 minutes (3440 seconds) Clip duration range 5.0s to 26.5s Mean clip duration 8.7s Median clip duration 7.5s… See the full description on the dataset page: https://huggingface.co/datasets/ayush712145/sarvam-tts-dataset.tabulartext-to-speechn<1K0 likes18 downloads3mo agoHugging Face19Khaledtelbahnasy /egyptian-arabic-stt-data Egyptian Arabic STT Dataset Synthetic Egyptian Arabic speech dataset generated by the Synthetic Egyptian Speech Data Pipeline. Samples are human-reviewed and quality-validated using Whisper ASR (WER/CER). Dataset Statistics Metric Value Total samples 50 Total duration 85.2s (0.02h) Dialect validated 50 / 50 Average WER 0.4136 Average CER 0.1642 Topics food_ordering Fields Field Type Description id string Deterministic… See the full description on the dataset page: https://huggingface.co/datasets/Khaledtelbahnasy/egyptian-arabic-stt-data.tabularautomatic-speech-recognitionn<1K1 likes16 downloads4mo agoHugging Face20arvinsingh /welsh-speech-dataset Welsh Speech Dataset A multimodal dataset of 33 speakers producing 10 Welsh phrases, captured using 3DMD technology with audio and dense facial landmark annotations. Dataset Overview Speakers: 33 participants Phrases: 10 Welsh phrases per speaker Sequences: ~330 (33 speakers x 10 phrases) Modalities: Audio recordings (.wav) 3D facial reconstructions (.obj meshes + texture maps) 68-point facial landmarks (ibug68 template) Fluency Scores: Each phrase rated 0-5 (5 =… See the full description on the dataset page: https://huggingface.co/datasets/arvinsingh/welsh-speech-dataset.tabularautomatic-speech-recognitionn<1K0 likes14 downloads8mo agoHugging Face21PalakEngineerMaster /Processed_TTS_Multilingual_Data Processed TTS Multilingual Data Validated and quality-checked multilingual speech datasets for TTS training, covering 12+ Indian languages. Datasets Included Subset Samples Hours Description indic_voices_r 239,684 548.8h Indic Voices_R — IVR recordings rasa 201,509 361.2h RASA — read speech (wiki, conv, book, news) indictts_iitm 155,236 253.6h Indic TTS (IIT Madras) — studio TTS recordings at 48kHz Total 596,429 1,163.6h Languages… See the full description on the dataset page: https://huggingface.co/datasets/PalakEngineerMaster/Processed_TTS_Multilingual_Data.tabulartext-to-speech100K<n<1M0 likes6 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.