CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.2k downloads7mo agoHugging Face02rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes533 downloads6mo agoHugging Face03uriel /audio_data_kaggle_train_taskaaudio100K<n<1M0 likes499 downloads1y agoHugging Face04uriel /audio_data_kaggle_train_taskb_audio100K<n<1M0 likes496 downloads1y agoHugging Face05uriel /audio_data_kaggle_train_taskbaudio100K<n<1M0 likes362 downloads1y agoHugging Face06MBZUAI /NADI-2025-Sub-task-3-allFor training and developing your models in the closed track, we provide the following datasets, which are publicly available on Hugging Face: The datasets represent a wide range of Arabic varieties and recording conditions, with over 85K training sentences in total. The datasets consist of dialectal, modern standard, classical, and code-switched Arabic speech and transcriptions. All except the Mixat and ArzEn subset are diacritized. Dataset Type Diacritized Train Dev MDASPC… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/NADI-2025-Sub-task-3-all.audio10K<n<100K1 likes298 downloads1y agoHugging Face07uriel /audio_data_kaggle_train_ne_taskcaudio100K<n<1M0 likes298 downloads1y agoHugging Face08kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes298 downloads7mo agoHugging Face09nadi-task2 /nadi2026-adi20-micro-25pct-knnvc NADI 2026 ADI20-micro — kNN-VC augmented (4 target voices) Voice-converted copy of the 25% stratified subset (seed 42) of UBC-NLP/NADI_2026_ADI20_micro, made with kNN-VC following Abdullah et al. 2025. Configs: voice_01–voice_04, 16,757 rows each, train split only. Validation/test audio is deliberately left natural. Target voices: 4 Arabic speakers from Common Voice (~60s each), gender-balanced, the same set used across all dialects. column meaning audio converted… See the full description on the dataset page: https://huggingface.co/datasets/nadi-task2/nadi2026-adi20-micro-25pct-knnvc.audio10K<n<100K0 likes249 downloads2mo agoHugging Face10HTill /dcase2025_task2_dev DCASE 2025 Task 2 - Development Dataset Attributes d1v, d2v, d3v are encoded as ClassLabels. Usage from datasets import load_dataset dataset = load_dataset('HTill/dcase2025_task2_dev', trust_remote_code=True) audioaudio-classification1K<n<10K2 likes148 downloads8mo agoHugging Face11M3-SLU /M3-SLU-Task2audio1K<n<10K0 likes131 downloads4mo agoHugging Face12Codec-SUPERB /dcase2016_task2_synth Dataset Card for "dcase2016_task2_synth" More Information needed audio1K<n<10K0 likes92 downloads3y agoHugging Face13uriel /audio_data_kaggle_test_taskb_audio1K<n<10K0 likes60 downloads1y agoHugging Face14M3-SLU /M3-SLU-Task1audio1K<n<10K0 likes51 downloads9mo agoHugging Face15jb1999 /spoken-nlp-tasks-24k Spoken Text Benchmarks for Audio LLM Evaluation TTS-synthesized audio versions of standard NLP text benchmarks, designed for evaluating audio/speech LLMs on tasks where the ground-truth text is known. These datasets were originally text-only; this resource provides spoken audio renditions so that audio LLMs can be evaluated on the same tasks and compared against text-only baselines. Dataset Description This dataset contains 24,000 WAV files: 1,000 utterances x 6 TTS… See the full description on the dataset page: https://huggingface.co/datasets/jb1999/spoken-nlp-tasks-24k.audioaudio-classification10K<n<100K0 likes48 downloads7mo agoHugging Face16uriel /audio_data_kaggle_test_taskcaudio1K<n<10K0 likes47 downloads1y agoHugging Face17SPRINGLab /asr-task-dataaudio1K<n<10K1 likes42 downloads2y agoHugging Face18renumics /dcase24_task10_loc1 DCASE 2024 Challenge Task 10 Development Dataset: Acoustic-based Traffic Monitoring - Location 1 subset Citation Bondi, L., Ghaffarzadegan, S., Damiano, S., Kumar, A., Wu, H.-H., Lin, W.-C., Das, S., Horst, H.-G., & Waterschoot, T. van . (2024). DCASE 2024 Challenge Task 10 Development Dataset: Acoustic-based Traffic Monitoring [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10700792 License Creative Commons Attribution-NonCommercial-ShareAlike 4.0… See the full description on the dataset page: https://huggingface.co/datasets/renumics/dcase24_task10_loc1.audio1K<n<10K0 likes41 downloads2y agoHugging Face19hamees /asr-taskaudio1K<n<10K0 likes28 downloads2y agoHugging Face20hamees /asr-task-testaudio1K<n<10K0 likes15 downloads2y agoHugging Face21uriel /audio_data_kaggle_test_taskbaudio1K<n<10K0 likes12 downloads1y agoHugging Face22amanuelbyte /try-task-translation-african-amh-enaudio10K<n<100K0 likes12 downloads7mo agoHugging Face23anime-sh /DCase2016_Task2_Hear_2021audion<1K0 likes10 downloads2y agoHugging Face24hamees /asr-task-hindiaudio1K<n<10K1 likes9 downloads2y agoHugging Face25uriel /audio_data_kaggle_test_taskaaudio1K<n<10K0 likes6 downloads1y agoHugging Face26M3-SLU /M3-SLU-Task2-sample 🎧 M3-SLU Task 2 — Sample Dataset 🗣️ Multi-Speaker, Multi-Turn, Multi-Modal Spoken Language Understanding 🌍 Overview The M3-SLU (Task 2 Sample) dataset is part of the M3-SLU Benchmark designed to evaluate speaker-attributed reasoning in multi-speaker, multi-turn conversations.It pairs long-form audio, transcripts, and contextual metadata, enabling fair comparison between cascade (SD + ASR + LLM) and end-to-end MLLMs. 👉 This sample includes 100 instances across 4… See the full description on the dataset page: https://huggingface.co/datasets/M3-SLU/M3-SLU-Task2-sample.audion<1K0 likes5 downloads11mo agoHugging Face27QuantumDeepak /voice_clone_taskaudion<1K0 likes3 downloads1y agoHugging Face28M3-SLU /M3-SLU-Task1-sample 🎧 M3-SLU Task 1 — Sample Dataset 🗣️ Multi-Speaker, Multi-Turn, Multi-Modal Spoken Language Understanding 🌍 Overview The M3-SLU (Task 1 Sample) dataset is part of the M3-SLU Benchmark designed to evaluate speaker-attributed reasoning in multi-speaker, multi-turn conversations.It pairs long-form audio, transcripts, and contextual metadata, enabling fair comparison between cascade (SD + ASR + LLM) and end-to-end MLLMs. 👉 This sample includes 100 instances across 4… See the full description on the dataset page: https://huggingface.co/datasets/M3-SLU/M3-SLU-Task1-sample.audion<1K0 likes3 downloads11mo agoHugging Face2927Group /shared_taskgatedaudio10K<n<100K0 likes1 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.