datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tasksThis dataset is for storing assets for https://huggingface.co/tasks and https://github.com/huggingface/huggingface.js/tree/main/packages/tasks
Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.DCASE2026-Task5-DevSet
DCASE 2026 Task 5 Audio-Dependent Question Answering (ADQA) Development Set
This is the official Development Set for DCASE 2026 Challenge Task 5: Audio-Dependent Question Answering (ADQA).
The ADQA task focuses on addressing "Textual Hallucination" in Large Audio-Language Models (LALMs) — where models pass audio understanding benchmarks by relying on text prompts and internal linguistic priors rather than actual audio perception. ADQA introduces a rigorous evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Harland/DCASE2026-Task5-DevSet.audio_data_kaggle_train_taskaaudio_data_kaggle_train_taskb_Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.osworld_tasks_filesaudio_data_kaggle_train_taskbCreative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.task-page-imagesaudio_data_kaggle_train_ne_taskcTaskmaster-OmniVox-9K-Vault-Preview
Taskmaster OmniVox - Verified Speech Data Evaluation Hub
Buy the proof. Scale the data.
Most speech-data purchases force teams to assess volume claims, technical quality, and licensing risk at the same time. OmniVox changes the order: inspect the evidence first, run a scoped pilot second, and expand only after the subset passes your technical and legal review.
This repository is the public evaluation hub for Taskmaster's multilingual speech catalog. It contains preview samples… See the full description on the dataset page: https://huggingface.co/datasets/TaskMasterVirtualAlly/Taskmaster-OmniVox-9K-Vault-Preview.nadi2026-adi20-micro-25pct-knnvc
NADI 2026 ADI20-micro — kNN-VC augmented (4 target voices)
Voice-converted copy of the 25% stratified subset (seed 42) of
UBC-NLP/NADI_2026_ADI20_micro, made with kNN-VC
following Abdullah et al. 2025.
Configs: voice_01–voice_04, 16,757 rows each, train split only.
Validation/test audio is deliberately left natural.
Target voices: 4 Arabic speakers from Common Voice (~60s each), gender-balanced,
the same set used across all dialects.
column
meaning
audio
converted… See the full description on the dataset page: https://huggingface.co/datasets/nadi-task2/nadi2026-adi20-micro-25pct-knnvc.NADI-2025-Sub-task-3-allFor training and developing your models in the closed track, we provide the following datasets, which are publicly available on Hugging Face: The datasets represent a wide range of Arabic varieties and recording conditions, with over 85K training sentences in total. The datasets consist of dialectal, modern standard, classical, and code-switched Arabic speech and transcriptions. All except the Mixat and ArzEn subset are diacritized.
Dataset
Type
Diacritized
Train
Dev
MDASPC… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/NADI-2025-Sub-task-3-all.otoSpeech-full-duplex-task-oriented-20h
Dataset Viewer
https://cc-task-oriented-preview.vercel.app/
Task Walkthrough
https://www.oto.earth/research/task-oriented-dataset.html
What each of the seven tasks is for, what the two speakers could each see, and
how the interaction log lines up with the audio.
otoSpeech-full-duplex-task-oriented-20h
Contact
This sample dataset is provided for research purposes. We maintain larger and
more diverse datasets.
For collaborations… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-task-oriented-20h.dcase2025_task2_dev
DCASE 2025 Task 2 - Development Dataset
Attributes d1v, d2v, d3v are encoded as ClassLabels.
Usage
from datasets import load_dataset
dataset = load_dataset('HTill/dcase2025_task2_dev', trust_remote_code=True)
M3-SLU-Task2dcase2016_task2_synth
Dataset Card for "dcase2016_task2_synth"
More Information needed
osworld_tasks_filesM3-SLU-Task1audio_data_kaggle_test_taskb_spoken-nlp-tasks-24k
Spoken Text Benchmarks for Audio LLM Evaluation
TTS-synthesized audio versions of standard NLP text benchmarks, designed for
evaluating audio/speech LLMs on tasks where the ground-truth text is known.
These datasets were originally text-only; this resource provides spoken audio
renditions so that audio LLMs can be evaluated on the same tasks and compared
against text-only baselines.
Dataset Description
This dataset contains 24,000 WAV files: 1,000 utterances x 6 TTS… See the full description on the dataset page: https://huggingface.co/datasets/jb1999/spoken-nlp-tasks-24k.asr-task-datadcase24_task10_loc1
DCASE 2024 Challenge Task 10 Development Dataset: Acoustic-based Traffic Monitoring - Location 1 subset
Citation
Bondi, L., Ghaffarzadegan, S., Damiano, S., Kumar, A., Wu, H.-H., Lin, W.-C., Das, S., Horst, H.-G., & Waterschoot, T. van . (2024). DCASE 2024 Challenge Task 10 Development Dataset: Acoustic-based Traffic Monitoring [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10700792
License
Creative Commons Attribution-NonCommercial-ShareAlike 4.0… See the full description on the dataset page: https://huggingface.co/datasets/renumics/dcase24_task10_loc1.audio_data_kaggle_test_taskcosworld_tasks_filesOS_World_Task_Navosworld_tasks_filesosworld_tasks_files
