CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ropedia-ai /xperience-10mgated ⚠️ Important: If you have already submitted an access request but have not completed the required DocuSign agreement, your request will remain pending. Please complete signing and we will grant access once verified. Interactive Intelligence from Human Xperience Xperience-10M Dataset Summary Xperience-10M is a large-scale egocentric multimodal dataset of human experience for embodied AI, robotics, world models, and spatial… See the full description on the dataset page: https://huggingface.co/datasets/ropedia-ai/xperience-10m.3dvideo-classification1M<n<10M249 likes84k downloads5mo agoHugging Face02japanese-asr /whisper_transcriptions.reazon_speech_allaudio10M<n<100M16 likes47k downloads2y agoHugging Face03ml-resources /Daimon-Infinity Daimon-Infinity mirror This repository is a file-preserving mirror of daimonrobotics/Daimon-Infinity on ModelScope. Source and license Upstream: daimonrobotics/Daimon-Infinity License: CC BY-NC-SA 4.0 Attribution: Daimon Robotics / Daimon-Infinity This mirror keeps the upstream directory layout and is distributed under the same CC BY-NC-SA 4.0 license. No data is altered; files are transferred with integrity checks supplied by ModelScope and the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/ml-resources/Daimon-Infinity.audio1K<n<10K2 likes38k downloads21d agoHugging Face04zaibihassan /Quranic-Recitation-Data 🌟 Overview Quranic Recitation Dataset (Word-by-Word Sync) is a highly optimized, production-ready dataset containing high-quality audio recitations of the Holy Quran synchronized at the word-by-word level. This dataset features 135 world-renowned reciters, with every Surah (114 chapters) mapped precisely to millisecond-accurate word timestamps. It is designed for modern Islamic mobile and web applications — served via a Cloudflare Edge CDN with native… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Recitation-Data.audioautomatic-speech-recognition10K<n<100K5 likes22k downloads16h agoHugging Face05RidheshBhati /Complete_Data_Source_100K_HOURS Multi-Language Audio Collection (100K Hours) This repository is physically reorganized for Absolute 100% Data Visibility. 🏗️ Global Consolidator Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here. audio1M<n<10M4 likes17k downloads5mo agoHugging Face06davidscripka /MIT_environmental_impulse_responsesMIT Environmental Impulse Response Dataset The audio recordings in this dataset are originally created by the Computational Audition Lab at MIT. The source of the data can be found at: https://mcdermottlab.mit.edu/Reverb/IR_Survey.html. The audio files in the dataset have been resampled to a sampling rate of 16 kHz. This resampling was done to reduce the size of the dataset while making it more suitable for various tasks, including data augmentation. The dataset consists of 271 audio files… See the full description on the dataset page: https://huggingface.co/datasets/davidscripka/MIT_environmental_impulse_responses.audioaudio-classificationn<1K9 likes16k downloads3y agoHugging Face07mshah1 /speech_robust_bench Dataset Card for "speech_robust_bench" More Information needed audio1M<n<10M11 likes14k downloads2y agoHugging Face08Jinxing1 /MQ-RAVSBench MQ-RAVSBench MQ-RAVSBench is a benchmark for mask-quality auditing in referring audio-visual segmentation. Each example links a video clip, audio, a referring expression, the ground-truth object mask, and candidate masks with different error patterns. The benchmark is used by MQ-Auditor to assess whether a candidate mask should be accepted, revised, or rejected. All paths stored in the metadata files are relative to the dataset root. Dataset Layout MQ-RAVSBench/… See the full description on the dataset page: https://huggingface.co/datasets/Jinxing1/MQ-RAVSBench.audio1K<n<10K2 likes12k downloads4mo agoHugging Face09pollen-robotics /reachy-mini-emotions-library Reachy Mini Emotions Library Curated emotion recordings for the Reachy Mini robot, maintained by Pollen Robotics. Each move is a JSON trajectory (head pose, antennas, body yaw, sampled over time) paired with an Opus audio track. Motion is sampled at 50 Hz; audio is mono Ogg/Opus (decoded natively by the robot). Requires reachy_mini ≥ v1.8.4 (its move loader resolves non-.wav audio sidecars). File layout Files live at the root of the dataset, named <emotion>.json +… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-emotions-library.audioroboticsn<1K18 likes12k downloads3mo agoHugging Face10NCSpeech /YO-CPT-ru YO-CPT-ru YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.audiotext-to-speech1M<n<10M16 likes10k downloads2mo agoHugging Face11ai4bharat /Rasagated Rasa: Towards Building an Expressive Multilingual Text-To-Speech Dataset for Indian Languages Funded by: Bhashini, Ministry of Electronics and Information Technology, Government of IndiaSupported by: EkStep Foundation and Nilekani Philanthropies Overview We introduce Rasa, the first high-quality multilingual expressive Text-to-Speech (TTS) dataset for any Indian language. It comprises a minimum of 20 hours per speaker with a target of covering a female and male… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Rasa.audiotext-to-speech1M<n<10M58 likes9.9k downloads4mo agoHugging Face12KRAFTON /Raon-OpenTTS-Pool Raon-OpenTTS-Pool Technical Report Raon-OpenTTS-Pool is a large-scale open English speech corpus for text-to-speech (TTS) training, constructed from 8 publicly available speech corpora and a set of web-sourced recordings. It is the training data behind Raon-OpenTTS, an open TTS model that performs on par with state-of-the-art closed-data systems. 615K hours of speech audio 239.7M speech segments 11 source datasets aggregated into a unified format All… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Pool.texttext-to-speech100M<n<1B43 likes9.8k downloads4mo agoHugging Face13Rocky131 /OmniReasoner-SFT OmniReasoner-SFT OmniReasoner-SFT is a mixed-source, research-only supervised fine-tuning dataset for audio-visual and long-video reasoning. It contains two-stage cold-start SFT trajectories with interval selection, zoom-in evidence, and final answers. Contents data/train.jsonl: HF-ready training JSONL with repo-relative media paths. media/: raw and derived media referenced by train.jsonl. manifests/media_manifest.jsonl: media inventory with repo paths, source family… See the full description on the dataset page: https://huggingface.co/datasets/Rocky131/OmniReasoner-SFT.audiovisual-question-answering10K<n<100K0 likes8.5k downloads4mo agoHugging Face14mythicinfinity /libritts_r Dataset Card for LibriTTS-R LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus (http://www.openslr.org/60/) which is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, published in 2019. Overview This is the LibriTTS-R dataset, adapted for the datasets library. Usage Splits There are 7 splits (dots replace dashes from the original dataset, to comply with… See the full description on the dataset page: https://huggingface.co/datasets/mythicinfinity/libritts_r.audiotext-to-speech100K<n<1M47 likes8.2k downloads3y agoHugging Face15ai4bharat /indicvoices_rgated IndicVoices-R: Multilingual, Multi-Speaker Speech Corpus for Indian TTS Dataset Summary IndicVoices-R (IV-R) is the largest multilingual Indian text-to-speech (TTS) dataset derived from an automatic speech recognition (ASR) dataset. It contains 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. This dataset is designed to enhance the development of robust Indian TTS models by providing diverse speaker demographics, natural… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indicvoices_r.audiotext-to-speech100K<n<1M44 likes7.8k downloads2y agoHugging Face16InteliLab /ict_s2s_refactoredaudio100K<n<1M0 likes7.5k downloads4mo agoHugging Face17japanese-asr /whisper_transcriptions.reazonspeech.allaudio10M<n<100M4 likes7.3k downloads2y agoHugging Face18FaisaI /tadabur-align-references tadabur-align-references Precomputed reference embeddings powering tadabur-align — word-level timestamp extraction for Quranic recitation via DTW alignment transfer (no ASR). What this is For 5,481 of the Quran's 6,236 ayahs, this dataset holds frame-level tadabur-embedding features for up to 8 reference reciters, plus each reference's word-level timestamps and internal-pause intervals. No audio is included — only model outputs and timing data. tadabur-align… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur-align-references.audio0 likes7.1k downloads2mo agoHugging Face19RidheshBhati /STT_MODEL Multilingual STT Dataset Audio and transcript pairs for 50 languages. Each language is a Dataset Viewer configuration with train, validation, and test splits. Language configurations amharic: Amharic arabic_msa: Arabic MSA assamese: Assamese bengali: Bengali czech: Czech dutch: Dutch egyptian_arabic: Egyptian Arabic english: English farsi_persian: Farsi - Persian filipino_tagalog: Filipino - Tagalog french: French german: German greek: Greek gujarati: Gujarati… See the full description on the dataset page: https://huggingface.co/datasets/RidheshBhati/STT_MODEL.audio1M<n<10M1 likes6.9k downloads1mo agoHugging Face20Reza2kn /telegram-audiobook-chizzled Telegram Persian Audiobook Chizzled 1,555,434 Persian audiobook clips · 14,430.384 hours · 16 kHz mono PCM WAV · public Parquet release This is a large, provenance-preserving collection of Persian audiobook audio gathered from 26 Telegram channels accessible to the collector account. Each source message is retained as message-level provenance and segmented with Silero voice-activity detection (VAD) into pause-aware clips. The audio bytes are embedded in Parquet files, so the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telegram-audiobook-chizzled.audion<1K4 likes6.7k downloads13d agoHugging Face21KBLab /rixvox-v2 RixVox-v2: A Swedish parliamentary speech dataset RixVox-v2 is a parliamentary speech dataset spanning nearly 23000 hours of speech. The dataset was built by matching and force aligning speeches in parliamentary protocols to media recordings of debates. Each observation contains metadata about the speaker's name, gender, district, role, party affiliation, and the date the speech was given. We include identifiers for protocols, speeches and speakers that allow linking observations in… See the full description on the dataset page: https://huggingface.co/datasets/KBLab/rixvox-v2.audioautomatic-speech-recognition1M<n<10M12 likes6.3k downloads1y agoHugging Face22RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K2 likes6k downloads3mo agoHugging Face23yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M52 likes5.9k downloads7mo agoHugging Face24RidheshBhati /Codemixed_New Codemixed ASR Dataset Unified collection of code-mixed ASR datasets. audioautomatic-speech-recognition100K<n<1M2 likes5.6k downloads5mo agoHugging Face25japanese-asr /ja_asr.reazon_speech_allaudio10M<n<100M7 likes5.3k downloads2y agoHugging Face26r3hab /training-movies-test-stuffaudion<1K0 likes5.3k downloads2y agoHugging Face27RUC-NLPIR /Omnimodal-Agent-SFT-2K OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.audioquestion-answering1K<n<10K9 likes4.7k downloads7mo agoHugging Face28parler-tts /libritts_r_filtered Dataset Card for Filtered LibriTTS-R This is a filtered version of LibriTTS-R. It has been filtered based on two sources: LibriTTS-R paper [1], which lists samples for which speech restoration have failed LibriTTS-P [2] list of excluded speakers for which multiple speakers have been detected. LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_filtered.audiotext-to-speech100K<n<1M24 likes4.7k downloads2y agoHugging Face29Vyvo-Research /Emilia-YODAS-ENaudio10M<n<100M4 likes4.3k downloads11mo agoHugging Face30SALT-Research /DeepDialogue-orpheus DeepDialogue-orpheus DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text. 🚨 Important Notice This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.audioaudio-classification100K<n<1M8 likes4.3k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.