CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lighthouse-emnlp2024 /Clotho-Moment Clotho-Moment This repository provides wav files used in Language-based Audio Moment Retrieval. Each sample includes long audio containing some audio events with the temporal and textual annotation. Project page: https://h-munakata.github.io/Language-based-Audio-Moment-Retrieval/ Code: https://github.com/line/lighthouse Split Train train/train-{000..715}.tar 37930 audio samples Valid valid/valid-{000..108}.tar 5741 audio samples Test test/test-{000..142}.tar 7569… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/Clotho-Moment.audioaudio-text-to-text10K<n<100K2 likes3.6k downloads8mo agoHugging Face02TTS-AGI /majestrino-unified-detailed-captions Majestrino Unified Detailed Captions Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption. Stats 4,658,407 samples 932 tar files (~1.1 GB each) ~1,017 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption caption_type — always unified_detailed_caption transcription — speech transcription (when available, normalized from multiple source keys) duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.audioaudio-classification1M<n<10M3 likes2.1k downloads6mo agoHugging Face03laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M15 likes2.1k downloads11mo agoHugging Face04sarulab-speech /commonvoice22_sidongated CV22-Sidon Overview This dataset hosts a release of Mozilla Common Voice 22 restored with the Sidon speech restoration model. Source: Mozilla Common Voice 22.0 Processing: Sidon denoising (sarulab-speech/sidon-v0.1) with 21 s chunks and 48 kHz reconstruction Format: WebDataset shards (.tar.gz) Manifest: paths.yaml enumerates every shard path for Hugging Face–style loading License: Original Common Voice license (CC0 1.0) Languages 137 language folders are… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon.audiotext-to-speech10M<n<100M30 likes1.9k downloads1y agoHugging Face05TTS-AGI /commonvoice22-sidon-dacvae CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents Source sarulab-speech/commonvoice22_sidon Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.audioautomatic-speech-recognition1M<n<10M1 likes1.1k downloads6mo agoHugging Face06ArnieRamesh /CounterStrike-1K-360-wds CounterStrike-1K — 360p WebDataset shards This repo contains the 360p shards of CounterStrike-1K. Use the main repo to browse the manifest, schema, and subsets. 360p is the recommended resolution for most training pipelines — the actions/state/events/metadata sidecars are identical to the 720p shards, so you can swap resolutions without touching downstream code. Quickstart Start a fresh uv project and add the loader: mkdir cs1k-demo && cd cs1k-demo uv init uv add… See the full description on the dataset page: https://huggingface.co/datasets/ArnieRamesh/CounterStrike-1K-360-wds.textvideo-classification10K<n<100K0 likes969 downloads5mo agoHugging Face07litagin /reazon-speech-v2-clonegated Reazon Speech v2 dataset mirror Original Dataset Source Hugging Face Dataset Page: reazon-research/reazonspeech Project Page: Reazon Research License This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction: TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/litagin/reazon-speech-v2-clone.audioautomatic-speech-recognition10K<n<100K12 likes822 downloads2y agoHugging Face08AlienKevin /sbs_cantonese SBS Cantonese Speech Corpus This speech corpus contains 435 hours of SBS Cantonese podcasts from Auguest 2022 to October 2023. There are 2,519 episodes and each episode is split into segments that are at most 10 seconds long. In total, there are 189,216 segments in this corpus. Here is a breakdown on the categories of episodes present in this dataset: Category SBS Channels Episodes news 中文新聞, 新聞簡報 622 business 寰宇金融 148 vaccine 疫苗快報 71 gardening 園藝趣談 58 tech 科技世界… See the full description on the dataset page: https://huggingface.co/datasets/AlienKevin/sbs_cantonese.audio10K<n<100K7 likes789 downloads3y agoHugging Face09mkrausio /audiosnippets-cleaned Dataset Summary This dataset is a processed version of mitermix/audiosnippets. The dataset contains audio snippets that have been cleaned and resampled, making it suitable for tasks like audio captioning, audio classification, or other audio-based machine learning applications. Processing Details Transcriptions and broken characters were removed. All MP3 audio files were resampled to 16kHz for consistency. The accompanying JSON metadata was made consistent. Entries with… See the full description on the dataset page: https://huggingface.co/datasets/mkrausio/audiosnippets-cleaned.audio1M<n<10M3 likes760 downloads2y agoHugging Face10ClareNie /Light-Omni-Training Light-Omni Training Dataset This repository contains the training data used by Light-Omni, a multimodal agent framework for reflexive video understanding with long-term memory. Light-Omni uses memory-augmented multimodal streams to train adapters for memory construction, response generation, and reaction/action control. Links Project page: https://clare-nie.github.io/Light-Omni/ Code: https://github.com/Clare-Nie/Light-Omni Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.audiovisual-question-answering100K<n<1M3 likes707 downloads3mo agoHugging Face11laion /common-voice-subset-for-clapaudion<1K1 likes528 downloads9mo agoHugging Face12CASIA-LM /OpenS2S_Datasets How to Use? Download, merge the files, and extract You can run the following command to merge the compressed file parts after downloading. cat en_response_wav.tar.gz.* > en_response_wav.tar.gz cat zh_response_wav.tar.gz.* > zh_response_wav.tar.gz audio100K<n<1M8 likes461 downloads1y agoHugging Face13OrcinusOrca /YouTube-Cantonese Cantonese Audio Dataset from YouTube This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models. Data Source and Processing The data was obtained through the following process: Download: Audio (.m4a) and available Cantonese subtitles (.srt for zh-TW, zh-HK, zh-Hant)… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-Cantonese.audioautomatic-speech-recognition100K<n<1M5 likes439 downloads1y agoHugging Face14adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-hokkien Taiwan-Tongues-ASR-CE-dataset-hokkien 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hokkien.audioautomatic-speech-recognition10K<n<100K7 likes422 downloads9mo agoHugging Face15cmeraki /audiofolder_webdatasetaudio100K<n<1M0 likes384 downloads2y agoHugging Face16recursal /reprocessed_singapore_national_speech_corpus Dataset Card for Reprocessed National Speech Corpus NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here. Dataset Details Dataset Description Dataset Description: The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.audiotext-generation1M<n<10M7 likes362 downloads2y agoHugging Face17mispeech /MECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks 📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF) Dataset Description MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks: Audio Captioning: Generating textual descriptions for given audio Audio Question Answering: Answering questions about given audio Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.audioaudio-classification10K<n<100K4 likes354 downloads5mo agoHugging Face18collabora /hindi-asr-wdsaudio1M<n<10M0 likes352 downloads1y agoHugging Face19KBlueLeaf /laion-coco-13m-taraudio10M<n<100M1 likes342 downloads1y agoHugging Face20BAAI /ChildMandaringated ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5 Introduction ChildMandarin is a comprehensive, open-source Mandarin Chinese speech dataset specifically designed for research on young children aged 3 to 5. This dataset addresses the critical lack of publicly available resources for this age group, enabling advancements in automatic speech recognition (ASR), speaker verification (SV), and other related fields. The dataset is released… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/ChildMandarin.audioautomatic-speech-recognition10K<n<100K46 likes269 downloads1y agoHugging Face21laion /timbre-audio-caption-pairsaudio100K<n<1M2 likes212 downloads9mo agoHugging Face22adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-zhtw Taiwan-Tongues-ASR-CE-dataset-zhtw 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。 🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-zhtw.audioautomatic-speech-recognition100K<n<1M2 likes208 downloads9mo agoHugging Face23laion /freesound-commercially-permissive-subset-with-captionsaudio100K<n<1M2 likes204 downloads11mo agoHugging Face24seungheondoh /cmd-audio-dumpaudio10K<n<100K1 likes191 downloads1y agoHugging Face25falabrasil /cetucaudioautomatic-speech-recognition100K<n<1M1 likes180 downloads1y agoHugging Face26Sh1man /common_voice_21_ru Dataset Description Набор данных validated.tsv отфильтрованный по down_votes = 0 📊 Статистика датасета Информация по сплитам 🔹 Тренировочный набор (train) Метрика Значение Количество семплов 93,531 Общая продолжительность 132.25 часов (476,089.70 секунд) Средняя продолжительность семпла 5.09 секунд 🔹 Валидационный набор (validate) Метрика Значение Количество семплов 38,836 Общая продолжительность 55.21… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/common_voice_21_ru.audio100K<n<1M5 likes170 downloads1y agoHugging Face27omni-chat /sharechatxaudio1M<n<10M2 likes167 downloads2y agoHugging Face28farsi-asr /ganjoor-chunked-asr-datasetaudio100K<n<1M2 likes153 downloads2y agoHugging Face29mitermix /audioset-with-grounded-captionsaudio1M<n<10M4 likes141 downloads1y agoHugging Face30hidude562 /operation-cappicinoaudio10K<n<100K0 likes132 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.