CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01carlosdanielhernandezmena /ravnursson_asr Dataset Card for ravnursson_asr Dataset Summary The corpus "RAVNURSSON FAROESE SPEECH AND TRANSCRIPTS" (or RAVNURSSON Corpus for short) is a collection of speech recordings with transcriptions intended for Automatic Speech Recognition (ASR) applications in the language that is spoken at the Faroe Islands (Faroese). It was curated at the Reykjavík University (RU) in 2022. The RAVNURSSON Corpus is an extract of the "Basic Language Resource Kit 1.0" (BLARK 1.0) [1] developed… See the full description on the dataset page: https://huggingface.co/datasets/carlosdanielhernandezmena/ravnursson_asr.audioautomatic-speech-recognition10K<n<100K3 likes869 downloads1y agoHugging Face02carlot /AIShell Dataset Card for "Aishell1" More Information needed audio100K<n<1M3 likes615 downloads3y agoHugging Face03hdong51 /Human-Animal-Cartoon Human-Animal-Cartoon dataset Our Human-Animal-Cartoon (HAC) dataset consists of seven actions (‘sleeping’, ‘watching tv’, ‘eating’, ‘drinking’, ‘swimming’, ‘running’, and ‘opening door’) performed by humans, animals, and cartoon figures, forming three different domains. We collect 3381 video clips from the internet with around 1000 for each domain and provide three modalities in our dataset: video, audio, and pre-computed optical flow. The dataset can be used for Multi-modal Domain… See the full description on the dataset page: https://huggingface.co/datasets/hdong51/Human-Animal-Cartoon.audiozero-shot-classification1K<n<10K5 likes544 downloads5mo agoHugging Face04AKP20 /carva-audio-libraryaudio10K<n<100K0 likes382 downloads17d agoHugging Face05anonymous-submission-dataset-1 /CaReCoS CaReCoS A medical acoustic question-answering dataset for reasoning over mel spectrograms of heart, lung, and cough sounds. Each record provides a clinical question, the mel-spectrogram image of a recording, a ground-truth answer, and the recording's clinical metadata. The task is purely visual: a model receives the spectrogram image together with the question and must reason over the spectrogram to produce the answer. The raw audio is not used as model input - the original .wav… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-dataset-1/CaReCoS.audioquestion-answeringn<1K0 likes286 downloads13d agoHugging Face06CarolinePascal /plug_socket_mixedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 50, "total_frames": 15095, "total_tasks": 1, "total_videos": 150, "total_audio": 150, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_mixed.audiorobotics10K<n<100K0 likes227 downloads1y agoHugging Face07Angeriod /in_car_commands_26 Dataset Card for "in_car_commands_26" More Information needed audio10K<n<100K0 likes173 downloads2y agoHugging Face08CarolinePascal /plug_socket_single_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 30, "total_frames": 11536, "total_tasks": 1, "total_videos": 90, "total_audio": 90, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_single_2.audiorobotics10K<n<100K0 likes166 downloads1y agoHugging Face09CarolinePascal /plug_socket_movedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 44, "total_frames": 15772, "total_tasks": 1, "total_videos": 132, "total_audio": 132, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:44" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_moved.audiorobotics10K<n<100K0 likes154 downloads1y agoHugging Face10cara-ai /VoiceTrace-BenchVoiceTrace-Bench VoiceTrace is a benchmark and unified framework for who-said-what speech retrieval: given a natural-language query about a speaker's identity or what they said, retrieve the matching audio document. Unlike conventional speaker verification or diarization benchmarks, VoiceTrace evaluates retrieval jointly over who is speaking and what is being said, across both single-speaker and multi-speaker conversational recordings. This repository hosts the VoiceTrace-Bench… See the full description on the dataset page: https://huggingface.co/datasets/cara-ai/VoiceTrace-Bench.audioaudio-classification1K<n<10K0 likes148 downloads3d agoHugging Face11Caruso77 /figli-e-napule-mediaaudio1K<n<10K0 likes111 downloads3d agoHugging Face12carlosdanielhernandezmena /chm150_asr Dataset Card for chm150_asr Dataset Summary The CHM150 is a corpus of microphone speech of mexican Spanish taken from 75 male speakers and 75 female speakers in a noise environment of a "quiet office" with a total duration of 1.63 hours. Speakers were encouraged to respond between some pre selected open questions or they could also describe a particular painting showed to them in a computer monitor. By so, the speech is completely spontaneous and one can see it in the… See the full description on the dataset page: https://huggingface.co/datasets/carlosdanielhernandezmena/chm150_asr.audio1K<n<10K0 likes108 downloads2y agoHugging Face13CarolinePascal /plug_socket_singleThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 65, "total_frames": 24999, "total_tasks": 1, "total_videos": 195, "total_audio": 195, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:65" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_single.audiorobotics10K<n<100K0 likes83 downloads1y agoHugging Face14Data-Science-Nigeria /voice-of-care-health-dataset Voice of Care AI for Global Health Benchmark Dataset Overview This dataset contains spoken Hausa Health datasets with rich annotations covering emotion, intent, speaker demographics, and dialect variation, intended for speech and NLP research. Dataset Summary Property Details Language Hausa Modality Audio + Text Task(s) e.g. Speech Recognition, Emotion Detection, Dialect Identification Version 1.0.0 🛠️ Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Data-Science-Nigeria/voice-of-care-health-dataset.audioautomatic-speech-recognition10K<n<100K1 likes82 downloads2mo agoHugging Face15kawshikbuet17 /bengali-telecom-customer-care-speech-v2 Bengali Telecom Customer Care Synthetic Speech Dataset v2 Dataset Description This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts. The dataset is intended for experiments with: Bengali ASR/STT Bengali TTS Speech-to-text preprocessing Telecom/customer-care domain adaptation Synthetic speech research This is a second version of the Bengali Telecom Customer Care Synthetic Speech Dataset. It follows the same… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech-v2.audiotext-to-speech1K<n<10K0 likes70 downloads3mo agoHugging Face16TheMindExpansionNetwork /void-carousel VOID CAROUSEL Dark psychedelic trance from graveyard orbit. Fictional entity VOID CAROUSEL is a fictional sentient derelict orbital carousel in the Sonic Forage universe: an abandoned amusement machine circling a dead world, translating its failing motors, empty passenger rings and intercepted signals into imagined psychedelic trance. It is not a human performer; its visual identity depicts no human likeness. This is an original fictional characterization, not a… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/void-carousel.audion<1K0 likes62 downloads6d agoHugging Face17carlosdanielhernandezmena /dummy_corpus_asr_esThis is an example of a repository where the audio files are not compressed in tar files. audion<1K0 likes58 downloads2y agoHugging Face18zachz /Human-Animal-Cartoon-PC-VAaudio1K<n<10K0 likes57 downloads5mo agoHugging Face19Angeriod /in_car_commands_60 Dataset Card for "in_car_commands_60" More Information needed audio10K<n<100K0 likes56 downloads2y agoHugging Face20CarolinePascal /plug_socket_single_slowThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 50, "total_frames": 24939, "total_tasks": 1, "total_videos": 150, "total_audio": 150, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_single_slow.audiorobotics10K<n<100K0 likes56 downloads1y agoHugging Face21jaeyong2 /cartesia-sonic-preview-ztts1-zero-shot-sample Cartesia Sonic on ZTTS1 zero-shot — sample with reference audio 100 utterances per language (700 rows) from the zero-shot subsets of ZTTS1-Eval, synthesized with Cartesia Sonic (preview) in voice-cloning mode. Unlike the full set, every row carries the reference recording as well as the synthesized clip, so a take can be compared against the voice it was cloning without checking out the benchmark. Columns column meaning audio the clip the model produced… See the full description on the dataset page: https://huggingface.co/datasets/jaeyong2/cartesia-sonic-preview-ztts1-zero-shot-sample.audiotext-to-speechn<1K0 likes50 downloads22d agoHugging Face22rgsgs /asoul_carol 声音数据 数据来源为asoul的珈乐 22年5月~21年6月的大部分录播时长共5小时 无内容标记已完成响度匹配数据在carol_fast_lzma2.zip里压缩算法是fast lzma2 太旧的解压软件可能不支持字母s开头的音频是歌声数据,量少质量低,建议删除无授权,侵删 2025.2备注 都2025年了还能每月五十个下载,都是神人了💧 audion<1K1 likes48 downloads2y agoHugging Face23CarolinePascal /plug_socket_single_paddingThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 50, "total_frames": 22752, "total_tasks": 1, "total_videos": 150, "total_audio": 150, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_single_padding.audiorobotics10K<n<100K0 likes47 downloads1y agoHugging Face24Titung /car-crash-audio-cc Car Crash Audio (Creative Commons) 46 clips (up to 40s each, varied length) sourced from 38 YouTube videos licensed Creative Commons -- Attribution, found via search for "car crash audio". Attribution CC BY requires attribution on reuse. Full per-clip attribution (title, channel, source URL) is in metadata.csv. Summary of unique source videos: Car Crash Sound Effect | Realistic Impact Audio for Films, Shorts & Game Development -- Stock Media (CC BY) Car Crash… See the full description on the dataset page: https://huggingface.co/datasets/Titung/car-crash-audio-cc.audioaudio-classificationn<1K0 likes44 downloads3mo agoHugging Face25carlicode /violence_contextaudion<1K1 likes39 downloads3y agoHugging Face26carankt /SmartHearingAids-data Semantic Hearing This repository provides code for the binaural target sound extraction model proposed in the paper, Semantic Hearing: Programming Acoustic Scenes with Binaural Hearables, presented at UIST'23. This model helps us create systems that let you control what you want to hear in the environment, in real-time, using noise-cancelling earbuds & headphones. https://github.com/vb000/SemanticHearing/assets/16723254/f1b33d8c-179a-4d50-92aa-6a99dde696d0 Conda environment… See the full description on the dataset page: https://huggingface.co/datasets/carankt/SmartHearingAids-data.audio0 likes36 downloads4mo agoHugging Face27sarayusapa /carnatic-ragasaudio1K<n<10K0 likes35 downloads6mo agoHugging Face28carlosdanielhernandezmena /dimex100_light Dataset Card for dimex100_light Dataset Summary The DIMEx100 LIGHT Corpus (DL) is a reduced version of the DIMEx100 Corpus (D100). DL was created in 2016 by Carlos Daniel Hernández Mena, with the aim of facilitating the use of the DIMEx100 Corpus in various automatic speech recognition systems. The most important differences between DIMEx100 LIGHT and the original are: The DL only contains audio files and transcriptions, unlike the D100 which contains pronunciation… See the full description on the dataset page: https://huggingface.co/datasets/carlosdanielhernandezmena/dimex100_light.audio1K<n<10K0 likes32 downloads2y agoHugging Face29Katock /carl_johnson_voice_pack Carl Johnson Voice Pack/Dataset audio1K<n<10K3 likes28 downloads3y agoHugging Face30CarolinePascal /record_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so100_follower", "total_episodes": 1, "total_frames": 433, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "audio_files_size_in_mb": 100, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/record_test.audioroboticsn<1K0 likes28 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.