datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.SwitchLingua_audio
Dataset Card for SwitchLingua_text
🚀 News
[19/09/2025] SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset is accepted by NeurIPS 2025!
[30/05/2024] The manuscript can be found on arXiv.
Dataset Summary
SwitchLingua is a comprehensive multilingual and multicultural code-switching dataset designed to advance research in automatic speech recognition, natural language processing, and conversational AI. The… See the full description on the dataset page: https://huggingface.co/datasets/Shelton1013/SwitchLingua_audio.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/kanepi-1977/Agent-Reasoning-WebSearch-260K.jam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.zoengjyutgaai
張悦楷講古語音數據集
English
呢個係張悦楷講《三國演義》、《水滸傳》、《走進毛澤東的最後歲月》、《鹿鼎記》語音數據集。張悦楷係廣州最出名嘅講古佬 / 粵語説書藝人。佢從上世紀七十年代開始就喺廣東各個收音電台度講古,佢把聲係好多廣州人嘅共同回憶。本數據集收集嘅係佢最知名嘅四部作品。
數據集用途:
TTS(語音合成)訓練集
ASR(語音識別)訓練集或測試集
各種語言學、文學研究
直接聽嚟欣賞藝術!
TTS 效果演示:https://huggingface.co/spaces/laubonghaudoi/zoengjyutgaai_tts
説明
所有文本都根據 https://jyutping.org/blog/typo/ 同 https://jyutping.org/blog/particles/ 規範用字。
所有文本都使用全角標點,冇半角標點。
所有文本都用漢字轉寫,無阿拉伯數字無英文字母
所有音頻源都存放喺/source,為方便直接用作訓練數據,切分後嘅音頻都放喺 opus/
所有 opus 音頻皆為… See the full description on the dataset page: https://huggingface.co/datasets/1111xxx/zoengjyutgaai.MultiModalDataset
Dataset Card for MultiModal Dataset
Dataset Description
Dataset Summary
MultiModal Dataset is a curated collection of 85,000 samples spanning three modalities: text, images, and audio. It combines high-quality web content, image-caption pairs from COCO 2017, and audio samples from AudioSet to enable comprehensive multimodal model training and evaluation.
The dataset is organized into three subsets:
fineweb: 37,500 high-quality web text samples (>8… See the full description on the dataset page: https://huggingface.co/datasets/lv12/MultiModalDataset.Ficbook-Audio-Instruct-10K
Ficbook Audio Instruct 10K
Synthetic audio instruction dataset for training Russian audio-language models.
Contains ~10K samples of fiction text voiced with OpenAI TTS and paired with diverse instruction tasks.
Dataset Description
This dataset was created for training and evaluating audio-language models on Russian fiction content.
Each sample contains:
Audio: Fiction text voiced using OpenAI's gpt-4o-mini-tts model
Text: Original text from ficbook stories
Question:… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Ficbook-Audio-Instruct-10K.panta_instruct_multi_modal_v1
Panta Instruct Multi-Modal v1
Dataset d'instructions multimodal en français : chaque exemple associe une question
(texte + parole + pictogrammes) à une réponse (texte + pictogrammes).
Colonnes
Colonne
Type
Description
audio
Audio (24 kHz, mono)
Enregistrement de la question (text_input)
text_input
string
Question / instruction
text_output
string
Réponse
pictos_input
list[string]
Identifiants des pictogrammes de la question
pictos_output… See the full description on the dataset page: https://huggingface.co/datasets/audibeal74/panta_instruct_multi_modal_v1.ONOTE
ONOTE: Omnimodal Notation Objective Topology Examination
ONOTE is a large-scale, omnimodal benchmark designed to evaluate music understanding across three major notation systems: Western Staff, Jianpu (Numbered Notation), and Guitar Tablature.
📂 Dataset Structure
The dataset is organized into two primary sub-directories based on the notation and instrument type:
1. pitch_Jianpu_dataset (Staff & Numbered Notation)
This sub-folder focuses on Western staff and… See the full description on the dataset page: https://huggingface.co/datasets/Weisiqing123/ONOTE.colloqialized_prompt
Colloquialized Prompt Dataset
This repository contains prompt and audio variants derived from the 60
WildClawBench tasks, plus the reusable task template. It supports experiments
that compare written prompts, spoken-style rewrites, synthesized speech, raw
ASR transcripts, and normalized ASR transcripts.
Dataset layout
.
├── prompts/ # Instructions used by rewrite/normalization jobs
├── scripts/ # Reproducible data preparation… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/colloqialized_prompt.medreport_audio_204
MedReport - Audio Dataset
Dataset Description
This dataset contains medical report audio files with their transcriptions, formatted according to HuggingFace Audio Dataset specifications. It's suitable for training speech-to-text models and instruction-following models in the medical domain.
Dataset Structure
This dataset follows the official HuggingFace Audio Dataset format:
dataset/
└── train/
├── audio/
│ ├── 20240315143022.wav
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/wouk1805/medreport_audio_204.hindi_voice_transcriptionsTheresaaudio-function-calling
Audio Function Calling Dataset
A synthetic multi-turn conversation dataset for audio-based tool/function calling.
Each sample contains a system prompt with tool definitions, alternating user and assistant turns,
where user turns are designed for audio (natural spoken language with filler words) and assistant
turns may include tool calls.
Note: This is the first batch (137 audio samples, 153 text-only samples). We are actively improving the generation pipeline and will be adding… See the full description on the dataset page: https://huggingface.co/datasets/nandinireddy123/audio-function-calling.11537606_ShenWanhong
