CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SALT-Research /DeepDialogue-orpheus DeepDialogue-orpheus DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text. 🚨 Important Notice This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.audioaudio-classification100K<n<1M8 likes2.3k downloads1y agoHugging Face02soynade-research /Bambara-Speech-Translation-Data AfVoices-Translated (Bambara-English) This is a Bambara speech translation dataset, which is built on the African Next Voices (AfVoices) Bambara ASR corpus. It provides English translations for the human-corrected subset of the original collection, creating a parallel corpus for Bambara-English machine translation and speech-to-text tasks. Methodology We machine-translated the human-validated transcriptions from AfVoices using the Oolel-translator repository. Inference… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Bambara-Speech-Translation-Data.audioautomatic-speech-recognition100K<n<1M1 likes951 downloads7mo agoHugging Face03reazon-research /reazonspeechgated Dataset Card for ReazonSpeech Dataset Summary This dataset contains a diverse set of natural Japanese speech, collected from terrestrial television streams. It contains more than 35000 hours of audio. Paper: ReazonSpeech: A Free and Massive Corpus for Japanese ASR Disclaimer TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF JAPANESE COPYRIGHT ACT ARTICLE 30-4. Dataset Format Audio files are available in FLAC… See the full description on the dataset page: https://huggingface.co/datasets/reazon-research/reazonspeech.automatic-speech-recognition10M<n<100M123 likes581 downloads2y agoHugging Face04SALT-Research /DeepDialogue-xtts DeepDialogue-xtts DeepDialogue-xtts is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the XTTS-v2 variant of the dataset, where speech is generated using XTTS-v2 with explicit emotional conditioning. 🚨 Important This dataset is large (~180GB) due to the inclusion of high-quality audio files. When cloning the… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-xtts.audioaudio-classification100K<n<1M8 likes195 downloads1y agoHugging Face05soynade-research /Wolof-ASR-DataA curated Wolof ASR dataset from various sources: Split Fleurs Alfa CV Kallama UB Total Train 8.72 16.13 34.97 33.60 4.52 97.94 Test 1.75 2.84 6.21 5.91 1.12 17.83 This dataset was used to finetune Wolof-HuBERT-Base for ASR. audioautomatic-speech-recognition10K<n<100K2 likes136 downloads7mo agoHugging Face06kadirnar /ai-researcher-roadmap-media AI Researcher Roadmap Media Optional video and subtitle assets for the AI Researcher Roadmap application. Repository layout manifest.json: file sizes and SHA-256 checksums used by the application. videos/<stem>.mp4: lecture video. subs/<stem>.<language>.vtt: subtitle tracks. subs/<stem>.asr.<language>.vtt: ASR-generated subtitle tracks. The application downloads only the selected lecture and its subtitle tracks. Files are cached locally and can be played offline… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/ai-researcher-roadmap-media.videoautomatic-speech-recognitionn<1K0 likes135 downloads2mo agoHugging Face07KSE-RESEARCH-Group /ukr-dialects-audio-dataset Ukrainian Dialects Audio Dataset Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits. Dataset Description This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets: NaUKMA-Audio-Dataset Ivanna-Stefiuk-Audio-Dataset Larysa-Irodenko-Audio-Dataset Hutsulendia-Audio-Dataset Dido-Yvanchyk-Audio-Dataset-v2 Dataset Structure train: 27,675 samples validation: 3… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukr-dialects-audio-dataset.audioautomatic-speech-recognition10K<n<100K1 likes89 downloads7mo agoHugging Face08researchaudio /apple-speechanalyzer-vs-whisper-cpp-mac Apple SpeechAnalyzer vs whisper.cpp on Mac Four complete speech-recognition benchmark runs over the same deterministic 40-speaker LibriSpeech test-clean snapshot: Engine Model path WER CER Repeated median post-speech latency Repeated p95 Apple SpeechAnalyzer progressiveTranscription on macOS 26.5 1.98% 1.02% 125–132 ms 194–201 ms whisper.cpp server 1.8.4 · ggml-small.en 4.28% 1.79% 122–125 ms 152–161 ms Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.tabularautomatic-speech-recognitionn<1K0 likes55 downloads2mo agoHugging Face09lilgoose777 /thaha-research-data2-v2gated Nepali Speech Dataset (YouTube-sourced) 441 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split. Splits train: 441 segments validation: 0 segments test: 0 segments Transcript columns — read this before training Each segment carries three transcript variants. They are NOT interchangeable: text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2-v2.audioautomatic-speech-recognitionn<1K0 likes49 downloads13d agoHugging Face10lilgoose777 /thaha-research-data1gated Nepali Speech Dataset (YouTube-sourced) 59 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split. Splits train: 59 segments validation: 0 segments test: 0 segments Transcript columns — read this before training Each segment carries three transcript variants. They are NOT interchangeable: text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data1.audioautomatic-speech-recognitionn<1K0 likes33 downloads25d agoHugging Face11lilgoose777 /thaha-research-data2gated Nepali Speech Dataset (YouTube-sourced) 63 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split. Splits train: 63 segments validation: 0 segments test: 0 segments Transcript columns — read this before training Each segment carries three transcript variants. They are NOT interchangeable: text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2.audioautomatic-speech-recognitionn<1K0 likes19 downloads24d agoHugging Face12lilgoose777 /thaha-research-data2-v2-v2gated Nepali Speech Dataset (YouTube-sourced) 76 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split. Splits train: 76 segments validation: 0 segments test: 0 segments Transcript columns — read this before training Each segment carries three transcript variants. They are NOT interchangeable: text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2-v2-v2.audioautomatic-speech-recognitionn<1K0 likes19 downloads24d agoHugging Face13AkAiNp /nepal-oral-speech-researchgated Nepal Oral Speech (Research) Gated research speech corpus for approved researchers and partner techs (AkAiNp). Clips with backbone license research only Includes self and guardian modes when research license was chosen Rich metadata for dialect / origin / consent versions Not a substitute for the private API vault (private clips never land here by default) Public users should use nepal-oral-packs-open and nepal-oral-demo instead. Access Request access on the… See the full description on the dataset page: https://huggingface.co/datasets/AkAiNp/nepal-oral-speech-research.automatic-speech-recognitionn<1K0 likes11 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.