CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01reach-vb /random-audiosaudion<1K2 likes745 downloads2y agoHugging Face02Ranjit /or_in_datasetaudioautomatic-speech-recognition10K<n<100K1 likes490 downloads3y agoHugging Face03ccmusic-database /timbre_range Dataset Card for Timbre and Range Dataset Dataset Summary The timbre dataset contains acapella singing audio of 9 singers, as well as cut single-note audio, totaling 775 clips (.wav format) The vocal range dataset includes several up and down chromatic scales audio clips of several vocals, as well as the cut single-note audio clips (.wav format). Supported Tasks and Leaderboards Audio classification Languages Chinese, English Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/timbre_range.audioaudio-classification1K<n<10K28 likes293 downloads7mo agoHugging Face04ahancock516 /oai-5g-srs-ranging-dataset OAI 5G NR SRS Ranging Captures Uplink Sounding Reference Signal (SRS) channel-estimate captures from a monolithic 5G NR software-defined-radio testbed, collected for SRS-based ranging experiments. The gNB (OpenAirInterface on a USRP X410, Band n78, 40 MHz / 106 PRB) configures each UE to transmit SRS; the gNB's per-SRS frequency-domain channel estimate, oversampled IDFT CIR, and ToA estimate are streamed off the PHY via OAI's T_tracer and recorded at a series of known… See the full description on the dataset page: https://huggingface.co/datasets/ahancock516/oai-5g-srs-ranging-dataset.audion<1K0 likes243 downloads2mo agoHugging Face05patrickvonplaten /random_imgaudion<1K0 likes192 downloads1y agoHugging Face06ahmfuad01 /random-chunks-dsaudio1K<n<10K0 likes182 downloads2mo agoHugging Face07Ransaka /SinhalaASR-testaudio10K<n<100K1 likes115 downloads3y agoHugging Face08dianavdavidson /indic_voices_hindi_only_plus_vaani_random_sample_34548_4616_seed_43_cleanaudio10K<n<100K0 likes100 downloads27d agoHugging Face09DeepFake-Audio-Rangers /Arabic_Audio_Deepfake ArAD Dataset (Arabic Audio DeepFake Dataset) Dataset SummaryThis dataset contains Arabic deepfake audio samples, focusing mainly on Levantine dialect with some examples in Standard Arabic. It was created using the RVC v2 framework, fine-tuned on a custom dataset of multi-dialect Arabic speech. The goal is to simulate real-world deepfake audio attacks by generating synthetic speech from actual recordings and voice messages. One of the first datasets to include real-world deepfake… See the full description on the dataset page: https://huggingface.co/datasets/DeepFake-Audio-Rangers/Arabic_Audio_Deepfake.audioaudio-to-audio10K<n<100K4 likes98 downloads2y agoHugging Face10Ransaka /SinhalaASR-1000audio1K<n<10K2 likes97 downloads3y agoHugging Face11dianavdavidson /indic_voices_hindi_only_random_sample_17274_2308_seed_42audio10K<n<100K0 likes97 downloads28d agoHugging Face12dianavdavidson /indic_voices_hindi_only_random_sample_17274_2308_seed_42_cleanaudio10K<n<100K0 likes88 downloads27d agoHugging Face13Rangasuthan /tamil-english-podcast-diarization Tamil-English Code-Mixed Podcast Diarization Dataset Dataset Summary This dataset contains long-form Tamil-English code-mixed podcast recordings annotated for speaker diarization research. The recordings consist of natural conversational speech with multiple speakers and realistic acoustic conditions, making the dataset suitable for evaluating diarization pipelines in real-world scenarios. The dataset is intended to support research in: Speaker diarization Code-mixed… See the full description on the dataset page: https://huggingface.co/datasets/Rangasuthan/tamil-english-podcast-diarization.audioautomatic-speech-recognitionn<1K1 likes86 downloads7mo agoHugging Face14dianavdavidson /indic_voices_hindi_only_random_sample_17334_2312_seed_42audio10K<n<100K0 likes80 downloads28d agoHugging Face15dianavdavidson /indic_voices_hindi_only_plus_vaani_random_sample_17274_2308_seed_43_cleanaudio10K<n<100K0 likes66 downloads27d agoHugging Face16Raniahossam33 /Egyptian_TTS3RSaudio1K<n<10K2 likes64 downloads2y agoHugging Face17ranaRan689 /SaudiTalk SaudiTalk: A Multi-Source Dialectal Speech Dataset from Saudi Arabia Dataset Summary SaudiTalk is a curated and human-verified Arabic speech dataset covering three major Saudi dialects: Hijazi, Ha’il, and Southern. The dataset is constructed from publicly available social media content and is designed to support research in Automatic Speech Recognition (ASR), dialect identification, and Arabic speech processing. Key Features 3 Saudi dialects:… See the full description on the dataset page: https://huggingface.co/datasets/ranaRan689/SaudiTalk.audion<1K0 likes60 downloads12d agoHugging Face18RandomAi99 /lili-romanian-single-speaker-piper-cleanaudio10K<n<100K0 likes59 downloads6mo agoHugging Face19chukypedro /nedu_randyPeteraudion<1K0 likes53 downloads28d agoHugging Face20Ranjit /displace_dev_dataaudio10K<n<100K0 likes35 downloads3y agoHugging Face21marccgrau /sbbdata_snr_random Dataset Card for "sbbdata_snr_random" More Information needed audio1K<n<10K0 likes27 downloads4y agoHugging Face22AdamMalyshev /eval_bell_ring_put_tape_in_bin_random_init_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 12, "total_frames": 6383, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "audio_files_size_in_mb": 100, "fps": 30, "splits": { "train": "0:12" }, "data_path":… See the full description on the dataset page: https://huggingface.co/datasets/AdamMalyshev/eval_bell_ring_put_tape_in_bin_random_init_test.audiorobotics1K<n<10K0 likes25 downloads9mo agoHugging Face23Raniahossam33 /radasaudio1K<n<10K0 likes23 downloads2y agoHugging Face24kadirnar /random_dataaudion<1K0 likes19 downloads2y agoHugging Face25random-sequence /flock-demo-automatic-speech-recognition-sectionsaudion<1K0 likes18 downloads7mo agoHugging Face26Raniahossam33 /Badasaudio1K<n<10K0 likes17 downloads3y agoHugging Face27rs545837 /ranveer_audiodatasetaudion<1K0 likes17 downloads2y agoHugging Face28Cybrpgs /corpus5-proposed-no-overlap-random-1000gated Corpus5 proposed-rule unflagged random sample This manually gated dataset contains 1,000 uniformly randomly selected rows from the fixed 2,715,793-row prepared Corpus5 snapshot on mac02. A row is eligible only when both proposed checks are false: the full chunk interval does not intersect positive-duration diarization turns from two distinct speakers; and no speaker-change/no-change disagreement is detected at aligned adjacent words in transcript1 and transcript2 and propagated… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/corpus5-proposed-no-overlap-random-1000.audioautomatic-speech-recognition1K<n<10K0 likes17 downloads4d agoHugging Face29DynamicSuperb /L2EnglishFluency_speechocean762-Rankingaudion<1K0 likes16 downloads2y agoHugging Face30Raniahossam33 /GAFASaudio1K<n<10K0 likes15 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.