datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.Light-Omni-Training
Light-Omni Training Dataset
This repository contains the training data used by Light-Omni, a multimodal
agent framework for reflexive video understanding with long-term memory.
Light-Omni uses memory-augmented multimodal streams to train adapters for
memory construction, response generation, and reaction/action control.
Links
Project page: https://clare-nie.github.io/Light-Omni/
Code: https://github.com/Clare-Nie/Light-Omni
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.jam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.jam-actions-acoustic-v0
Dataset Card for jam-actions-acoustic-v0
Version: 1.0.2
Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI.
Summary
108 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence).
This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.Gurbani-MahanKosh-Frontier-Corpus
ੴ Gurbani & Bhai Kahn Singh Nabha Mahan Kosh Frontier Corpus
☬ ਗੁਰਬਾਣੀ ਅਤੇ ਭਾਈ ਕਾਹਨ ਸਿੰਘ ਨਾਭਾ 'ਮਹਾਨ ਕੋਸ਼' ਪ੍ਰਮਾਣਿਕ ਡਾਟਾਸੈੱਟ
👨💻 Project Lead & Architecture
Curator & Developer: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile: github.com/gurpreetsingh5523-source
Flagship Project: AMRIT Research OS (Autonomous Medical AI)
📖 Dataset Overview
An authoritative lexical dataset compiling authentic definitions… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Gurbani-MahanKosh-Frontier-Corpus.lolaby-traces
Lolaby — generation traces
Pipeline traces from Lolaby, an AI-powered lullaby generator built for the Build Small Hackathon 2026 (Backyard AI track).
Each trace is a complete witness of one end-to-end generation: every input the user gave, every model that ran, every prompt and raw output, every timing measurement, and the final audio. Published under CC0 so anyone can study, replay, or remix the pipeline.
What's in a trace
Each subfolder is one generation. Files:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/lolaby-traces.universeset
UniVerseSet
The training split of UniVerse (同谣).Held-out evaluation lives in UniVerseBench.
UniVerseSet is the post-training corpus for large audio–language models on world folk music: ASR, captions, and audio-grounded chat, plus automatically transcribed ABC scores.
「诗言志,歌永言,声依永,律和声。」—《尚书·舜典》
Sister dataset (benchmark)
universe-team/universebench
Live museum demo
http://143.89.224.8:8790/
What's here
Archives (download and unpack;… See the full description on the dataset page: https://huggingface.co/datasets/universe-team/universeset.colloqialized_prompt
Colloquialized Prompt Dataset
This repository contains prompt and audio variants derived from the 60
WildClawBench tasks, plus the reusable task template. It supports experiments
that compare written prompts, spoken-style rewrites, synthesized speech, raw
ASR transcripts, and normalized ASR transcripts.
Dataset layout
.
├── prompts/ # Instructions used by rewrite/normalization jobs
├── scripts/ # Reproducible data preparation… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/colloqialized_prompt.balanced-emotion-dataset-majestrino-withtemporal-detailed-captions
Balanced Emotion Dataset — Majestrino with Temporal Detailed Captions
An emotion-balanced subset of TTS-AGI/majestrino-unified-detailed-captions-temporal.
Overview
Total samples: 482,594
Samples per emotion category: 12,997
Number of emotion categories: 40
Format: WebDataset (tar files with FLAC audio + JSON metadata)
Number of tar files: 483
Samples per tar: ~1000
Balancing Strategy
Samples were selected from the source dataset using keyword matching on… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/balanced-emotion-dataset-majestrino-withtemporal-detailed-captions.synthetic-patient-dr-data
Synthetic Patient DR Data
Synthetic doctor-patient consultation dataset with structured clinical outputs and optional full-consultation audio.
Dataset Summary
This dataset was generated for research and prototyping in:
clinical dialogue generation
structured clinical extraction
text-to-audio workflows
conversational healthcare modeling
All consultations are synthetic and should not be treated as real clinical encounters.
Export Metadata
Mode: audio
Repo… See the full description on the dataset page: https://huggingface.co/datasets/TumeloKonaite/synthetic-patient-dr-data.hindi_voice_transcriptionsTheresatxt2vst
txt2vst Dataset
10,000 natural-language-to-VST-spec pairs for music plugin generation.
Dataset Description
Each sample maps a natural language description of a VST instrument to a structured spec.json that defines the complete plugin architecture.
Fields
prompt: Natural language description (e.g., "drum machine with kick snare and acid bass, punchy mastering")
completion: Compact JSON spec defining the plugin (voices, FX, theme, mastering chain)… See the full description on the dataset page: https://huggingface.co/datasets/fabriziosalmi/txt2vst.testdspodcast-transcripts
Podcast Transcripts Dataset
This dataset contains transcripts from Bitcoin and cryptocurrency podcasts,
processed by the belief-engines ETL pipeline.
Files
transcripts.parquet - Full episode transcripts with speaker diarization
transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings
Schema
transcripts.parquet
episode_id: Unique episode identifier
podcast_slug: Podcast name slug
episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.tuantientisettw-daily-dialogue-audio
Dataset Card for tw-daily-dialogue-audio
本資料集是一份臺灣日常情境的對話腳本(dialogue scripts)資料集,每筆樣本包含對話分類、主題、文字內容以及說話者輪廓/場景/天氣等情境 metadata。可作為文字→語音(TTS)合成、對話 ASR 評測之素材設計來源。共 29,337 筆樣本。
Dataset Details
Dataset Description
資料以對話腳本(純文字)為主,搭配豐富的情境 metadata:
category:對話類別(如:餐廳、醫療、家庭、商業、交通等)。
theme:該段對話的細部主題。
text:對話腳本本文(可包含多輪、多角色)。
meta:情境 metadata,包含:
profile:說話者輪廓
loc:場景/地點
weather:當下天氣
可用於下游語音/對話應用:將腳本送入 TTS 管線生成多角色音訊、評測 dialogue-aware ASR、訓練具備情境感知的對話模型。… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-daily-dialogue-audio.tdtu-vi-audio
