CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rajjanardhan00 /Seamless_Dummy_Dataset_Fixed_3 MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes266 downloads1y agoHugging Face02Codyfederer /tr-full-dataset TR-Full_dataset This is a merged speech dataset containing 41427 audio segments from 88 source datasets. Dataset Information Total Segments: 41427 Speakers: 222 Languages: tr Emotions: neutral, angry, sad, happy Original Datasets: 88 Dataset Structure Each example contains: audio: Audio file (WAV format, original sampling rate preserved) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.audioautomatic-speech-recognition10K<n<100K6 likes145 downloads1y agoHugging Face03Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes138 downloads1mo agoHugging Face04rorosese /my-voxtral-datasetaudion<1K0 likes135 downloads1y agoHugging Face05rodenhhh /ContextTTS_dataset ContextTTS Evaluation Dataset This is the official evaluation dataset for the paper "[ContextTTS Eval: A Benchmark for Evaluating Long-Form Contextual Expressive Text-to-Speech]". It is designed to evaluate the performance of multi-modal speech synthesis, specifically focusing on context-aware prosody and timbre consistency in Chinese conversations and audiobooks. Dataset Summary The dataset consists of high-quality Chinese audio-text pairs, organized into three distinct… See the full description on the dataset page: https://huggingface.co/datasets/rodenhhh/ContextTTS_dataset.audiotext-to-speechn<1K0 likes113 downloads6mo agoHugging Face06ZHANGYUXUAN-zR /MCF-Dataset MCF: Text LLMS For Multimodal Emotional Causality Data Dataset task definition and annotation example of the MCF framework. The framework contains two core subtasks: five-tuple element extraction (identifying Target, Holder, Aspect, Opinion, Sentiment, and Rationale) and sentiment chain analysis (constructing causal relationship chains between emotional events). The dataset is provided with the following structure. Each sample includes video, audio, and dialogue… See the full description on the dataset page: https://huggingface.co/datasets/ZHANGYUXUAN-zR/MCF-Dataset.audiotext-classification10K<n<100K2 likes72 downloads1y agoHugging Face07David-A-Amoo /naijavoices_dataset_85_hours_tts_bestFull&nbsp;dataset&nbsp;re-upload&nbsp;with&nbsp;more&nbsp;statistics&nbsp;as&nbsp;well&nbsp;as&nbsp;filtering&nbsp;scripts&nbsp;that&nbsp;give&nbsp;the&nbsp;top&nbsp;x&nbsp;files&nbsp;or&nbsp;best&nbsp;x&nbsp;hours&nbsp;for&nbsp;tts&nbsp;based&nbsp;and&nbsp;calculations&nbsp;from&nbsp;acoustic&nbsp;metrics\color{Blue}{\large \textbf{Full dataset re-upload with more statistics as well as filtering scripts that give the top x files or best x hours for tts based and calculations from acoustic… See the full description on the dataset page: https://huggingface.co/datasets/David-A-Amoo/naijavoices_dataset_85_hours_tts_best.tabular10K<n<100K2 likes62 downloads2mo agoHugging Face08sunbv56 /song_dataset 🎵 Vietnamese Song Lyrics and Word Timestamps Dataset Dataset Summary The song_dataset provides high-quality Vietnamese song data, including metadata, full lyrics, and particularly word-level timestamps. This dataset is optimally designed for tasks such as: Training and evaluating automatic speech recognition (ASR) models on music. Lyrics synchronization (Lyrics Alignment / Karaoke generation). Natural language processing (NLP) analysis on song lyrics. The data is… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset.text1K<n<10K0 likes39 downloads7mo agoHugging Face09jeju-potato /jeju_potato_datasetsaudio10K<n<100K0 likes36 downloads1y agoHugging Face10electron-rare /mascarade-dsp-dataset Mascarade — DSP & Signal Processing Q&A ✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11) Per-sample Stack Exchange Electronics attribution recovered via the SE /search/advanced + /questions/{id} API search : 169 samples (~5.35 %) confirmed as Stack Exchange Electronics (CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution (URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60). 535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/mascarade-dsp-dataset.texttext-generation1K<n<10K0 likes30 downloads5mo agoHugging Face11mira-iitjmu /ns-urdu-datasetaudio1K<n<10K0 likes29 downloads5mo agoHugging Face12ajey6789 /bhashini-datasetaudio1K<n<10K0 likes28 downloads1y agoHugging Face13galammadin-asr /chlid-datasetaudio10K<n<100K0 likes26 downloads7mo agoHugging Face14chtugha /small-german-medical-dialogue-dataset-for-moshi Small german dialogue dataset This dataset contains 500 completely made up medical phonecall dialogues between patients and a GP's office. Dataset Details Dataset Description 500 made up phonecalls that were first created with AI as text. The audio was then created using Openai tts-1-hd and the accurately timestamped transcripts were added. The audio files are formatted like this: Stereo with split channels: Speaker A is on the left channel… See the full description on the dataset page: https://huggingface.co/datasets/chtugha/small-german-medical-dialogue-dataset-for-moshi.audioaudio-text-to-textn<1K0 likes24 downloads4mo agoHugging Face15sunbv56 /song_dataset_chunked Vietnamese Songs Word-Level Timestamp Dataset (Chunked) This dataset contains word-level timestamp information for Vietnamese songs, specifically pre-chunked into segments up to 30 seconds for use in training or fine-tuning speech recognition (ASR) systems like Whisper. Dataset Summary The song_dataset_chunked provides high-quality Vietnamese song data, properly segmented into optimal ~30-second sequences. Duration Insights: Train split (train_chunked.jsonl): ~ 230.62… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset_chunked.tabular10K<n<100K0 likes17 downloads7mo agoHugging Face16yadorigi /Onomatopoeia_Dataset🎧 Onomatopoeia Dataset (Audio → Manga Expression) 音声解析結果をもとに、日本語のオノマトペ(擬音語・擬態語)を生成するためのデータセットです。 本データセットは、音そのものではなく、音から推定された特徴・空間・情景を入力とする構造化データであり、 漫画的な表現生成を目的としたマルチモーダルデータです。 📌 Dataset Summary 本データセットは以下のパイプラインから生成されています: Audio ↓ Audio Features (04_features.json) ↓ Audio Events (05_audio_events.json) ↓ Space Judgement (06_space_judgement.json) ↓ Scene Interpretation (07_scene_interpretation.json) ↓ Onomatopoeia (08_onomatopoeia.json) 👉 音 → 空間 → 情景 → オノマトペ という段階的生成構造を持ちます。 📊… See the full description on the dataset page: https://huggingface.co/datasets/yadorigi/Onomatopoeia_Dataset.texttext-generationn<1K0 likes16 downloads6mo agoHugging Face17mszawerd /concept-datasetaudio1K<n<10K0 likes13 downloads2y agoHugging Face18Tnaot /SPS-Bopha-Voice-Dataset-v1gated VibeVoice Fine-Tuning Dataset: SPS-Bopha-Voice-Dataset-v1 This dataset is formatted for fine-tuning VibeVoice. Structure training_data.jsonl: The main manifest file containing transcriptions and paths. chunks_staging/: Directory containing the audio clips. Usage with VibeVoice Clone this repository: git clone https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1 cd SPS-Bopha-Voice-Dataset-v1 Run the training script pointing to… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1.audiotext-to-speech1K<n<10K0 likes13 downloads10mo agoHugging Face19anian0707 /scasr_datasetaudio10K<n<100K0 likes13 downloads1mo agoHugging Face20Ailiance-fr /mascarade-dsp-dataset Mascarade — DSP & Signal Processing Q&A ✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11) Per-sample Stack Exchange Electronics attribution recovered via the SE /search/advanced + /questions/{id} API search : 169 samples (~5.35 %) confirmed as Stack Exchange Electronics (CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution (URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60). 535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-dsp-dataset.texttext-generation1K<n<10K0 likes12 downloads5mo agoHugging Face21solbon1212 /fma-dataset-aug-caption FMA-CLAP Caption Augmentation Dataset Overview This dataset is an enhanced version of the FMA (Free Music Archive) dataset, where we have augmented the original metadata with natural language captions generated using the CLAP (Contrastive Language-Audio Pretraining) model. The captions describe the genre, style, mood, and instrumentation of each track, making it more suitable for zero-shot learning, music classification, and text-to-music generation tasks.… See the full description on the dataset page: https://huggingface.co/datasets/solbon1212/fma-dataset-aug-caption.textaudio-classification1K<n<10K1 likes11 downloads2y agoHugging Face22RidheshBhati /Indic_New_dataset_TTS Indic TTS Dataset Hub (Mozilla) Validated audio–text pairs for multiple Indic languages from Mozilla Common Voice. Select the language from the Subset dropdown in the Dataset Viewer. Columns audio: WAV audio clip (16kHz) text: transcription duration: length in seconds speaking_rate: characters per second audio10K<n<100K0 likes11 downloads7mo agoHugging Face23chihhab999 /topipl-spanish-datasetaudio1K<n<10K0 likes9 downloads9mo agoHugging Face24IbraahimLab /voice-dataset Voice Dataset Collected from the web uploader tool. Voice Dataset Collected from the web uploader tool. audion<1K0 likes9 downloads7mo agoHugging Face25vichetkao /khmer_speech_news_datasetgated Khmer Speech Dataset Processing This repository contains scripts and instructions for preparing a Khmer speech dataset for machine learning tasks, such as automatic speech recognition (ASR). It demonstrates how to process a collection of audio files and metadata, and save them as Parquet files for efficient use in your training pipelines—without needing torchcodec. All dataset audio and transcripts in this project are sourced from https://wmc.org.kh/, the official website of… See the full description on the dataset page: https://huggingface.co/datasets/vichetkao/khmer_speech_news_dataset.audioaudio-classification10K<n<100K1 likes6 downloads7mo agoHugging Face26seleznevchamp1 /chester_bennington_30tracks_datasetaudion<1K0 likes5 downloads7mo agoHugging Face27ThisUsernameAlreadyExistsAlreadyExists /aitf-dfk3-synthetic-audio-datasetgatedaudioaudio-classification1K<n<10K0 likes5 downloads6mo agoHugging Face28RaccoON8579 /indicvc-datasetaudio100K<n<1M0 likes4 downloads3mo agoHugging Face29Diomande /s2o-datasetaudio10K<n<100K0 likes1 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.