datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.zoengjyutgaai
張悦楷講古語音數據集
English
呢個係張悦楷講《三國演義》、《水滸傳》、《走進毛澤東的最後歲月》、《鹿鼎記》語音數據集。張悦楷係廣州最出名嘅講古佬 / 粵語説書藝人。佢從上世紀七十年代開始就喺廣東各個收音電台度講古,佢把聲係好多廣州人嘅共同回憶。本數據集收集嘅係佢最知名嘅四部作品。
數據集用途:
TTS(語音合成)訓練集
ASR(語音識別)訓練集或測試集
各種語言學、文學研究
直接聽嚟欣賞藝術!
TTS 效果演示:https://huggingface.co/spaces/laubonghaudoi/zoengjyutgaai_tts
説明
所有文本都根據 https://jyutping.org/blog/typo/ 同 https://jyutping.org/blog/particles/ 規範用字。
所有文本都使用全角標點,冇半角標點。
所有文本都用漢字轉寫,無阿拉伯數字無英文字母
所有音頻源都存放喺/source,為方便直接用作訓練數據,切分後嘅音頻都放喺 opus/
所有 opus 音頻皆為 48000… See the full description on the dataset page: https://huggingface.co/datasets/CanCLID/zoengjyutgaai.legco-speech
香港立法會會議語音數據集
本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。
數據集製作流程
先去香港特別行政區立法會 YouTube下載所有會議紀錄並轉為 16kHz 採樣率嘅 OPUS音頻
用 fsmn-vad 切分所有語音,並用 Qwen3-ASR-1.7B 轉寫成粵文 srt 字幕
轉寫後用正則表達式修正字幕中常見轉寫錯誤
將數據集分成 raw、 segmented 兩個子集傳到HF
子集 subset
raw
segment
總行數 Row number
14,036
9,557,109
總時長 Total duration
22,195.55 hr (79,903,980.00 s)
20471.21 hr (73,696,365.27 s)
平均時長 Average duration
1.58 hr (5692.79 s)
7.71… See the full description on the dataset page: https://huggingface.co/datasets/laubonghaudoi/legco-speech.Audio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.OmniVideoBench
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
✨ Overview
Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, most existing benchmarks fail to jointly evaluate both audio and visual reasoning — often focusing on one modality or overlooking their interaction.
🎬 OmniVideoBench fills this gap.It’s a large-scale, rigorously curated… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/OmniVideoBench.wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.Speech2Latex
Speech2Latex Dataset
The Speech2LaTeX dataset is the first fully open-source large-scale dataset for converting spoken mathematical expressions and sentences into LaTeX. It comprises over 66,000 human-annotated audio samples of mathematical equations and sentences in both English and Russian, drawn from diverse scientific domains. This work lays the groundwork for future advances in multimodal AI, with a particular focus on mathematical content recognition.
The dataset was presented… See the full description on the dataset page: https://huggingface.co/datasets/marsianin500/Speech2Latex.distilled-web
Dataset Card for lenamerkli/distilled-web
This dataset consists of web-scraped data using a custom crawler purpose-built for each website.
Dataset Details
Dataset Sources
Repository: https://github.com/lenamerkli/distilled-web
Uses
This dataset is useful for training large language models.
The train split provides instruction-following and chat data for supervised fine-tuning (SFT) and instruction tuning.
The pretrain split… See the full description on the dataset page: https://huggingface.co/datasets/lenamerkli/distilled-web.SwitchLingua_audio
Dataset Card for SwitchLingua_text
🚀 News
[19/09/2025] SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset is accepted by NeurIPS 2025!
[30/05/2024] The manuscript can be found on arXiv.
Dataset Summary
SwitchLingua is a comprehensive multilingual and multicultural code-switching dataset designed to advance research in automatic speech recognition, natural language processing, and conversational AI. The… See the full description on the dataset page: https://huggingface.co/datasets/Shelton1013/SwitchLingua_audio.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.MultimodalMathBenchmarks
MultimodalMathBenchmarks
This repository contains the datasets for the paper Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs (ACL Findings 2026).
It covers the public benchmark datasets and their modality assets (text, images, and audio) used to evaluate the arithmetic capabilities of multimodal LLMs.
Canonical Upload Manifest
HF path
Local source
Count
Purpose
SharedMultimodalGrid.csv
SavedData/SharedMultimodalGrid.csv… See the full description on the dataset page: https://huggingface.co/datasets/cjerzak/MultimodalMathBenchmarks.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.Light-Omni-Training
Light-Omni Training Dataset
This repository contains the training data used by Light-Omni, a multimodal
agent framework for reflexive video understanding with long-term memory.
Light-Omni uses memory-augmented multimodal streams to train adapters for
memory construction, response generation, and reaction/action control.
Links
Project page: https://clare-nie.github.io/Light-Omni/
Code: https://github.com/Clare-Nie/Light-Omni
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.Marco_Longspeech
Marco-LongSpeech Dataset
Marco-LongSpeech is a multi-task long speech understanding dataset containing 8 different speech understanding tasks designed to benchmark Large Language Models on lengthy audio inputs.
📊 Dataset Statistics
Task Statistics
Task
Train
Val
Test
Total
Unique Audios
ASR
71,275
15,273
15,274
101,822
101,822
Temporal_Relative_QA
5,886
1,261
1,262
8,409
8,409
summary
4,366
935
937
6,238
6,238… See the full description on the dataset page: https://huggingface.co/datasets/ATH-MaaS/Marco_Longspeech.VoiceAssistant-Eval
🔥 VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
[🌐 Homepage]
[🔮 Visualization]
[💻 Github]
[📖 Paper]
[📊 Leaderboard ]
[📊 Detailed Leaderboard ]
[📊 Roleplay Leaderboard ]
🚀 Data Usage
from datasets import load_dataset
for split in ['listening_general', 'listening_music', 'listening_sound', 'listening_speech',
'speaking_assistant', 'speaking_emotion', 'speaking_instruction_following'… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/VoiceAssistant-Eval.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.episodes
My Weird Prompts - Episode Dataset
The production record of every episode of the
My Weird Prompts podcast: metadata, the transcript,
and the generation telemetry for how each episode was made - model, pipeline
version, GPU, timings and compute cost.
5,333 episodes. Synced daily from the production database.
from datasets import load_dataset
ds = load_dataset("My-Weird-Prompts/episodes", split="train")
Which dataset do you want?
This one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/episodes.MyMentorLLM-dataset
Dataset Card for MyMentorLLM
This dataset accompanies the MyMentorLLM paper, arXiv:2607.25667, which describes the simulation environment, generation procedure, experimental design and validation analyses. If you use this dataset, please cite the paper (see Citation Information).
Dataset Summary
MyMentorLLM is a multimodal voice- and text-based deliberate-practice environment used to generate 2,100 simulated Cognitive Behavioural Therapy (CBT) training sessions… See the full description on the dataset page: https://huggingface.co/datasets/RodolfoRizzi/MyMentorLLM-dataset.anonymous-storybench
Omni-StoryBench
Omni-StoryBench is a context-aware omnimodal story generation benchmark.Each sample provides a current story page and requires generating the next page's image, narration text, and speech utterance.
Dataset Structure
The dataset contains:
data/testset.jsonl: Main benchmark file.
images/: Page images.
texts/: Page text files.
speech/: Generated speech audio files.
instruction/: Source-level instruction metadata.
Data Fields
Each JSONL sample… See the full description on the dataset page: https://huggingface.co/datasets/omnibench/anonymous-storybench.SpeechInstructBench
SpeechInstructBench
Arxiv: https://arxiv.org/abs/2503.02769
This is the SpeechInstructBench dataset download page.
SpeechInstructBench is a multilingual (Chinese and English) benchmark designed to evaluate the instruction-following capabilities of speech models. Instruction-following refers to a model’s ability to accurately interpret and execute user-provided natural language directives while strictly adhering to all specified constraints and requirements. To comprehensively… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/SpeechInstructBench.reprocessed_singapore_national_speech_corpus
Dataset Card for Reprocessed National Speech Corpus
NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here.
Dataset Details
Dataset Description
Dataset Description:
The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.sdf_dataset_en
SpeechDialogueFactory Dataset
Background
This dataset is part of the SpeechDialogueFactory project, a comprehensive framework for generating high-quality speech dialogues at scale. Speech dialogue datasets are essential for developing and evaluating Speech-LLMs, but existing datasets face limitations including high collection costs, privacy concerns, and lack of conversational authenticity. This dataset addresses these challenges by providing synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/minghanw/sdf_dataset_en.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.IndicCMix
IndicCMix
Most Indic NLP data assumes people write in one script and one language at a time. Real chat looks nothing like that. You get Hindi words in Roman letters, English verbs in the middle of a Tamil sentence, and the same person switching scripts halfway through a paragraph.
This dataset is an attempt to cover that actual messiness. For every English sentence, you get three different Indic renderings of it: one code-mixed in the native script, one clean native-script… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicCMix.strudel-rl-data
strudel-rl-data
Datasets from the Strudel-RL project (training a Qwen3.6-35B-A3B to write house / techno / hypnotic techno as Strudel programs).
sft/: SFT datasets v1..v5 (messages format, system+user+assistant) with stats.
synth/: GLM-5.3-Flash synthetic rounds 1..5 (prompt, code, reasoning length).
captions/: GLM captions of programs. prompts/: brief generators and held-out eval briefs (v0..v3; v3 = reference-driven).
gold/: 26 hand-written gold programs + manifest.… See the full description on the dataset page: https://huggingface.co/datasets/amol-derick/strudel-rl-data.OmniRewardBench
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
📄 Paper |
💻 Code |
🤗 Benchmark (This Dataset) |
🤗 Training Data |
🤗 Model |
🏠 Homepage
Reward models (RMs) play a critical role in aligning AI behaviors with human preferences, yet they face two fundamental challenges: (1) Modality Imbalance, where most RMs are mainly focused on text and image modalities, offering limited support for video, audio, and other modalities; and… See the full description on the dataset page: https://huggingface.co/datasets/HongbangYuan/OmniRewardBench.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.OpenDialog_English
OpenDialog English
This dataset contains English dialog and conversation data.
Dataset Structure
The dataset is provided in Parquet format with 153 splits for efficient loading.
Data Files
Format: Parquet
Splits: 153 files (train-00001-of-00153.parquet through train-00153-of-00153.parquet)
Total Size: ~72.8 GB
Loading the Dataset
from datasets import load_dataset
# Load the full dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/OpenDialog_English.
