datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zoengjyutgaai
張悦楷講古語音數據集
English
呢個係張悦楷講《三國演義》、《水滸傳》、《走進毛澤東的最後歲月》、《鹿鼎記》語音數據集。張悦楷係廣州最出名嘅講古佬 / 粵語説書藝人。佢從上世紀七十年代開始就喺廣東各個收音電台度講古,佢把聲係好多廣州人嘅共同回憶。本數據集收集嘅係佢最知名嘅四部作品。
數據集用途:
TTS(語音合成)訓練集
ASR(語音識別)訓練集或測試集
各種語言學、文學研究
直接聽嚟欣賞藝術!
TTS 效果演示:https://huggingface.co/spaces/laubonghaudoi/zoengjyutgaai_tts
説明
所有文本都根據 https://jyutping.org/blog/typo/ 同 https://jyutping.org/blog/particles/ 規範用字。
所有文本都使用全角標點,冇半角標點。
所有文本都用漢字轉寫,無阿拉伯數字無英文字母
所有音頻源都存放喺/source,為方便直接用作訓練數據,切分後嘅音頻都放喺 opus/
所有 opus 音頻皆為 48000… See the full description on the dataset page: https://huggingface.co/datasets/CanCLID/zoengjyutgaai.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.Light-Omni-Training
Light-Omni Training Dataset
This repository contains the training data used by Light-Omni, a multimodal
agent framework for reflexive video understanding with long-term memory.
Light-Omni uses memory-augmented multimodal streams to train adapters for
memory construction, response generation, and reaction/action control.
Links
Project page: https://clare-nie.github.io/Light-Omni/
Code: https://github.com/Clare-Nie/Light-Omni
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.MultimodalMathBenchmarks
MultimodalMathBenchmarks
This repository contains the datasets for the paper Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs (ACL Findings 2026).
It covers the public benchmark datasets and their modality assets (text, images, and audio) used to evaluate the arithmetic capabilities of multimodal LLMs.
Canonical Upload Manifest
HF path
Local source
Count
Purpose
SharedMultimodalGrid.csv
SavedData/SharedMultimodalGrid.csv… See the full description on the dataset page: https://huggingface.co/datasets/cjerzak/MultimodalMathBenchmarks.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.reprocessed_singapore_national_speech_corpus
Dataset Card for Reprocessed National Speech Corpus
NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here.
Dataset Details
Dataset Description
Dataset Description:
The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.ConceptCaps
Dataset Card for ConceptCaps
Dataset Summary
ConceptCaps is a music captioning dataset derived from MusicCaps, specifically designed for concept-based interpretability research in text-to-audio (TTA) generation systems. The dataset provides categorized musical concept annotations from a distilled taxonomy (200 unique tags) alongside natural language captions, enabling fine-grained analysis of how TTA models represent and generate musical concepts.
Unlike existing datasets… See the full description on the dataset page: https://huggingface.co/datasets/bsienkiewicz/ConceptCaps.audio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.music-codes
互联网歌曲数据集
使用encodec编码的歌曲数据,sampling_rate=24000, bandwidth=6.0, n_codebooks=8, 截取时长=30s
各数据列说明
id : 歌曲源id
name : 歌曲名
singer : 歌手名
text : 歌词
code : 原音乐使用EnCodec编码的结果, 是形状为[T,8]的二维列表。
clotho-jaQwen/Qwen3-4B-Instruct-2507を使用してClothoを日本語に翻訳したデータです。
ライセンスは元データに従います。
Gurbani-MahanKosh-Frontier-Corpus
ੴ Gurbani & Bhai Kahn Singh Nabha Mahan Kosh Frontier Corpus
☬ ਗੁਰਬਾਣੀ ਅਤੇ ਭਾਈ ਕਾਹਨ ਸਿੰਘ ਨਾਭਾ 'ਮਹਾਨ ਕੋਸ਼' ਪ੍ਰਮਾਣਿਕ ਡਾਟਾਸੈੱਟ
👨💻 Project Lead & Architecture
Curator & Developer: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile: github.com/gurpreetsingh5523-source
Flagship Project: AMRIT Research OS (Autonomous Medical AI)
📖 Dataset Overview
An authoritative lexical dataset compiling authentic definitions… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Gurbani-MahanKosh-Frontier-Corpus.conversation-bench
Conversation Bench
75-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a conference assistant for the AI Engineer World's Fair.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a conference assistant for the AI Engineer World's Fair, handling session registration, schedule queries, speaker lookups, and… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/conversation-bench.una-fraza-al-diya
Una fraza al diya
Ladino language learning sentences prepared by Karen Sarhon of Sephardic Center of Istanbul. Each sentence has translations in Turkish, English, Spanish. Includes audio and image. 307 sentences in total.
Source: https://sefarad.com.tr/judeo-espanyolladino/frazadeldia/
Citation
If you use this dataset, please cite:
Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish
Preparing an endangered language for the digital age: The… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/una-fraza-al-diya.AURA-Chat-Edit
AURA-Chat-Edit: Conversational Music Editing Dataset
Overview
AURA-Chat-Edit is a large-scale conversational music editing dataset used to train AURA, a unified multimodal framework for conversational music editing. The dataset contains 66,539 multi-turn dialogues pairing natural-language edit instructions with structured edit-token outputs across 7 edit types.
Each dialogue simulates a user requesting a music edit (e.g., "remove the drums", "add a jazzy… See the full description on the dataset page: https://huggingface.co/datasets/OpenRB-Lab/AURA-Chat-Edit.colloqialized_prompt
Colloquialized Prompt Dataset
This repository contains prompt and audio variants derived from the 60
WildClawBench tasks, plus the reusable task template. It supports experiments
that compare written prompts, spoken-style rewrites, synthesized speech, raw
ASR transcripts, and normalized ASR transcripts.
Dataset layout
.
├── prompts/ # Instructions used by rewrite/normalization jobs
├── scripts/ # Reproducible data preparation… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/colloqialized_prompt.balanced-emotion-dataset-majestrino-withtemporal-detailed-captions
Balanced Emotion Dataset — Majestrino with Temporal Detailed Captions
An emotion-balanced subset of TTS-AGI/majestrino-unified-detailed-captions-temporal.
Overview
Total samples: 482,594
Samples per emotion category: 12,997
Number of emotion categories: 40
Format: WebDataset (tar files with FLAC audio + JSON metadata)
Number of tar files: 483
Samples per tar: ~1000
Balancing Strategy
Samples were selected from the source dataset using keyword matching on… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/balanced-emotion-dataset-majestrino-withtemporal-detailed-captions.clotho-jaQwen/Qwen3-4B-Instruct-2507を使用してClothoを日本語に翻訳したデータです。
ライセンスは元データに従います。
audio-function-calling
Audio Function Calling Dataset
A synthetic multi-turn conversation dataset for audio-based tool/function calling.
Each sample contains a system prompt with tool definitions, alternating user and assistant turns,
where user turns are designed for audio (natural spoken language with filler words) and assistant
turns may include tool calls.
Note: This is the first batch (137 audio samples, 153 text-only samples). We are actively improving the generation pipeline and will be adding… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/audio-function-calling.ONOTE
ONOTE: Omnimodal Notation Objective Topology Examination
ONOTE is a large-scale, omnimodal benchmark designed to evaluate music understanding across three major notation systems: Western Staff, Jianpu (Numbered Notation), and Guitar Tablature.
📂 Dataset Structure
The dataset is organized into two primary sub-directories based on the notation and instrument type:
1. pitch_Jianpu_dataset (Staff & Numbered Notation)
This sub-folder focuses on Western staff and… See the full description on the dataset page: https://huggingface.co/datasets/chengxin666/ONOTE.EdgeMMEval
EdgeMMEval
Minimal multimodal evaluation dataset for on-device inference testing.
Covers functional correctness, accuracy, latency stress, and memory
pressure across image, audio, text, multi-turn, combination, structured
output, and tool-calling cases.
Dataset summary
The test split is defined in data/test/metadata.jsonl (200 rows). Each
row has a test_id (for example IMG-001, STO-020) and a modality.
Modality
Samples
Focus
Image
34
VQA, OCR, description… See the full description on the dataset page: https://huggingface.co/datasets/CortexSwarm/EdgeMMEval.acp-corpus-filtered
ACP Filtered Conversations (via Cortico)
This dataset is a filtered slice of conversation recordings and transcripts
from the American Conversation Project (ACP), retrieved via
Cortico's ACP integration. "Filtered" means every
fragment included here already passed an LLM-based salience pass (the
project's internal "wheat vs. chaff" filter) that removed filler, small talk,
and interjections — everything kept is a substantive, quote-anchored moment
someone actually said.
This is… See the full description on the dataset page: https://huggingface.co/datasets/omzugo/acp-corpus-filtered.audio-function-calling
Audio Function Calling Dataset
A synthetic multi-turn conversation dataset for audio-based tool/function calling.
Each sample contains a system prompt with tool definitions, alternating user and assistant turns,
where user turns are designed for audio (natural spoken language with filler words) and assistant
turns may include tool calls.
Note: This is the first batch (137 audio samples, 153 text-only samples). We are actively improving the generation pipeline and will be adding… See the full description on the dataset page: https://huggingface.co/datasets/nandinireddy123/audio-function-calling.11537606_ShenWanhong
